Proprietary Pipeline

Data Cleaning

Scanned archives, handwritten tables, blurry reports — into clean, structured data. Transparent performance, protected internals.

100%
table digits correct (120/120, cell-by-cell human audit)
98.3%
independent second engine (118/120) — cross-checked, not self-checked
0
known errors after merge across 5 page types
~1h
per 189-page report, a few RMB in cost

What it handles

Pure scanned documents with no text layer — the hardest shape of archive data. Typewriter text mixed with handwriting, hand-drawn tables, formulas, drawings, stamps and signatures.

Typewriter textnormalised automatically — legacy notation (I for 1, comma decimals) fixed 100%
Hand-drawn tablesdashed rules, rotated headers — read cell by cell into clean tables
Handwritten formulaschemistry & math — converted, with uncertain terms flagged for review
Geological profileslandscape drawings — embedded, with searchable captions
Stamps & signaturespreserved as-is, no hallucinated reconstruction
Blurry scansno aggressive pre-processing — we never invent strokes that are not there
Multiple languagesverified on Russian archives; per-language notation rules handled per archive type

How it works

Three steps, one handoff. You never lose control of your documents — the conversion result belongs to you.

1 · Send the scan
drop a sample PDF — typewriter pages, tables, drawings, anything. Get a sample conversion back.
2 · Automatic conversion + cross-check
every page is converted and cross-verified; anything uncertain is explicitly flagged — never guessed.
3 · Reviewed handoff
flagged items get a quick human pass, then the full structured file is delivered. Yours to keep.

Case studies

A 1975 geological exploration report — 189 pages, 300 DPI scan, every digit verified by hand against the original.

Chemical analysis table · 120 cells

120/120 correct
Scan (input)
Scanned chemical analysis table
Structured output
%15699591411285064620
SiO₂55.8268.5663.1153.6968.5057.5072.9065.6863.91
TiO₂0.900.450.800.800.550.790.480.570.59
Al₂O₃16.0915.6315.6318.7614.8017.3314.1816.2515.88
Fe₂O₃2.252.532.123.272.963.781.652.383.22

Oxide column sums re-verified against geochemistry rules automatically. The hardest cell — Al₂O₃ 18.76 — was missed twice by the second engine and caught by the primary.

Handwritten formula page

converted · uncertain terms flagged
Scan (input)
Scanned handwritten formula page
Structured output

Δg = g·α²/2

Handwritten symbols converted to real math; anything the engine is unsure about is explicitly marked for human review instead of guessed.

Landscape geological profile

embedded + searchable caption
Scan (input)
Scanned geological profile drawing
Structured output

Drawings stay as images inside the structured document, with a generated caption so they become searchable — no attempt to redraw what was drawn by hand.

How it was verified

300 DPI crops audited cell by cell by hand, on a real 1975 report — not a synthetic benchmark.

EngineTable digitsBody text
Primary engine100% (120/120)zero hallucination; uncertain words marked (?)
Independent second engine98.3% (118/120)1–3 rationalised errors per page — used only for cross-checking
Merged output0 known errorsacross all 5 demo page types

The pipeline also caught and corrected 7 pre-existing errors in the client's own manually-edited reference data — including a lost entire row (п.п.п) and a misread city name.

Cost & speed

Few RMB per reporta 189-page, 96.9 MB scanned PDF converts for roughly the price of a cup of coffee
~1 hour wall timewith concurrency; single-stream ≈ 1.5–2 min per page
1.5 MB self-contained HTMLformulas as real math, one file, no viewer needed
Talk to us about your archives →

anmuning@awareliquid.ai · bring a sample document, get a sample conversion