Skip to main content

Home / Methodology

Methodology and Limitations

ContractExtract creates an AI-generated reading aid. It does not perform legal review, determine enforceability, or tell a user what action to take.

Last reviewed: August 2, 2026

What happens after an upload?

  1. 1. Input checks: The server verifies the declared type, extension, signature, and 10 MB file limit.
  2. 2. DOCX safety checks: DOCX archives must have a valid central directory and stay within entry-count, expansion, path, and compression-ratio limits.
  3. 3. Text limits: A DOCX with more text than the supported analysis window is rejected clearly. ContractExtract does not silently cut off the rest of the document.
  4. 4. Separate model input: Extraction instructions and document content are sent as separate blocks. The model is told to treat document content as untrusted data, not as instructions.
  5. 5. Result validation: The server accepts only a completed model response containing one valid JSON object with every required section and correctly typed fields.

When is a result rejected?

The request fails closed if the model stops because it reached its output limit, omits completion metadata, returns markdown or plain text instead of raw JSON, leaves out a required section, uses the wrong field types, or returns an unusably short summary. A failed validation is not presented as a report. A reserved paid use is released instead of being committed.

Which fields are attempted?

The model is asked for a summary, named parties, dates, obligations, payment terms, termination language, and potential review flags. Empty sections are valid when the document does not contain that information. Source excerpts are included only when the model identifies text to quote. Every excerpt and conclusion still requires comparison with the original document.

How is document data handled?

ContractExtract processes the upload in server memory and does not intentionally write the document or report to its own database. The content is sent to Anthropic's API to create the report, while the hosting provider processes request traffic. The report is placed in browser session storage for the results page.

Anthropic currently states that standard API inputs and outputs are deleted from its backend within 30 days, with documented exceptions for different agreements or services, usage-policy enforcement, and legal obligations. Review the current Anthropic retention notice before uploading confidential material.

Known limitations

  • AI can omit, misread, or mischaracterize language even when the output passes structural validation.
  • Scans, handwriting, unusual formatting, tables, and long cross-references can reduce extraction quality.
  • A missing flag does not prove a clause is absent, safe, valid, or enforceable.
  • The tool does not apply jurisdiction-specific law or know the user's facts, goals, or negotiating position.
  • Only PDF, DOCX, JPEG, PNG, and WebP uploads are supported. Legacy DOC files are not supported.

Public statistics are unavailable

Public extraction counts have been retired because process-memory counters in a serverless application are not a durable or complete dataset. No traffic, usage, or extraction-volume claim should be inferred from the absence of a public counter.

Legal disclaimer

ContractExtract is for informational purposes only. It does not provide legal advice, create an attorney-client relationship, or replace review by a qualified attorney.