Message threading and near-duplicate detection
We built two analytics services for a client in the eDiscovery industry, and the platform that runs them. They find the repetition in large document sets, so legal teams review less.
- Collected documentsEmail and files gathered for a legal matter, often in the millions.
- Message threadingRebuilds each conversation so a reviewer reads it once.
- Near-duplicate detectionGroups documents that are almost the same and picks one to review first.
- ReviewLegal teams read far fewer documents to cover the same material.
The problem
In a legal matter, every document a reviewer reads costs time and money, and much of what is collected is repetition. The same email appears quoted inside every reply to it. The same contract appears in a dozen drafts that differ by a sentence.
The work is to find that repetition accurately, across very large document sets, and quickly enough to keep a review moving.
What we built
- Message threading. It reads each message, separates what the sender wrote from what they quoted, and reconstructs the conversation. It handles the header and date formats that different mail systems and languages produce.
- Near-duplicate detection. It fingerprints the text of each document, groups the ones that are nearly identical, and chooses a principal document for each group.
One analytics platform
The two services are not separate applications. Each is a module on a platform we built, and they share everything that is not specific to the analysis.
- Built to extend. The second service reused the foundations of the first, so it took a fraction of the effort a standalone application would have.
- Automated deployment. Building, testing, and releasing run from scripts kept with the code, the same way in every environment. Releasing is routine.
- Parallel by design. Each step of an analysis is its own service, so large document sets are processed in batches side by side. This is the source of the dramatic improvement in processing time.
Client details are confidential, so this page describes the work without naming them.