Datasets
Buyer room
TURBOLENS.IOManage supplier listing
TurboLens (product: DocumentLens) — Two distinct latent assets. (1) A labelled multilingual trade-document corpus: every production extraction is an image-to-structured-JSON pair on real SEA bills of lading, customs declarations, commercial invoices, packing lists, certificates of origin, phytosanitary certificates, delivery orders and proof-of-delivery scans — across Thai, Vietnamese, Bahasa Indonesia/Malay, Tagalog, Khmer, Chinese and Japanese, across 50+ carrier layouts, with field-level ground truth confirmed by downstream customs submission. (2) A visual-forensics corpus: forged/tampered…
supplier index read from their site
Two distinct latent assets. (1) A labelled multilingual trade-document corpus: every production extraction is an image-to-structured-JSON pair on real SEA bills of lading, customs declarations, commercial invoices, packing lists, certificates of origin, phytosanitary certificates, delivery orders and proof-of-delivery scans — across Thai, Vietnamese, Bahasa Indonesia/Malay, Tagalog, Khmer, Chinese and Japanese, across 50+ carrier layouts, with field-level ground truth confirmed by downstream customs submission. (2) A visual-forensics corpus: forged/tampered versus genuine document images with splicing, copy-move, retouching and AI-generated/edit labels at suspicious-region coordinate level, plus document version-pairs with semantic-diff annotations.
From 3 yearsCoverage Information TechnologyAsset class Equities
Sample, licence terms, pricing and eval results when Turbolens publishes them. Until then, discover alternatives today with a 7-day trial.
Start 7-day trial →What buyers ask about Turbolens, answered from this page.
TurboLens (product: DocumentLens) — Two distinct latent assets. (1) A labelled multilingual trade-document corpus: every production extraction is an image-to-structured-JSON pair on real SEA bills of lading, customs declarations, commercial invoices, packing lists, certificates of origin, phytosanitary certificates, delivery orders and proof-of-delivery scans — across Thai, Vietnamese, Bahasa Indonesia/Malay, Tagalog, Khmer, Chinese and Japanese, across 50+ carrier layouts, with field-level ground truth confirmed by downstream customs submission. (2) A visual-forensics corpus: forged/tampered…
Turbolens offers (Alternative, Supply Chain) — ~0.5-5 million processed document pages cumulatively (est. — freemium daily scan allowances across a low-thousands user base plus enterprise API throughput on shipment volumes, accumulated only since a ~2024 launch; the vendor publishes no processing totals, so this is an order-of-magnitude bound and should be sized against their actual logs in diligence).
Training and evaluating document-AI/OCR models on low-resource SEA scripts and real-world degraded scans (the scarce part — most public document corpora are English and clean); training tamper and synthetic-media detectors on real forged trade documents with region-level labels, which are almost absent publicly; fine-tuning models on customs-submission schema mapping for ASEAN Single Window; building document-version-diff and contract-redlining models from annotated pairs; and preference data where a human reviewer accepted or corrected an extraction.
The data is with 3 years of history.
Coverage spans APAC, US, Japan; Application Software; alternative, supply_chain; equities.
Turbolens has not added their own details yet. Not yet on file: