Complex Document Information Processing (CDIP) dataset

Published by National Institute of Standards and Technology | National Institute of Standards and Technology | Metadata Last Checked: June 27, 2025 | Last Modified: 1996-01-01 00:00:00

This dataset is called the "IIT CDIP collection". "CDIP" stands for "Complex Document Information Processing" and "IIT" stands for "Illinois Institute of Technology" who originally built the dataset. The dataset consists of documents from the states' lawsuit against the tobacco industry in the 1990s. As a result of the settlement of that lawsuit (the "Master Settlement Agreement"), the companies had to make all the documents public in an archive, which currently resides at UCSF, the University of California, San Francisco.IIT used this data to build a dataset of "messy" documents that were challenging for existing systems to process. There is handwriting on the documents, stains, etc. TREC used an automatic text conversion of this dataset in the TREC Legal Track, and we also have the original TIFF scans of the documents. The dataset consists of around 7 million documents, preprocessed with 90s-era OCR, and also the original page scans in TIFF format. See contact information in this record for access to this dataset.

Find Related Datasets

Click any tag below to search for similar datasets

Complete Metadata

bureauCode	[ "006:55" ]
identifier	ark:/88434/mds2-2531
landingPage	https://data.nist.gov/od/id/mds2-2531
language	[ "en" ]
programCode	[ "006:045" ]
theme	[ "Information Technology:Data and informatics" ]