Methodology

How we measure extraction quality

This page is about method, not a headline number. It describes how we test whether the extraction reads a drawing correctly, what has to be true before a change ships, and what we deliberately do not claim.

Why there is no percentage at the top of this page

You have probably seen a competitor publish a single accuracy figure. We do not, because a single figure is not a fact about your drawings. Extraction quality varies with the field being read, with whether the sheet is a CAD export or a fourth-generation photocopy, and with the drawing conventions of whoever drafted it. A number that averages all of that tells you what a vendor's sample looked like, not what your bid package will do.

What we can tell you honestly is how we measure, what we hold ourselves to, and what happens when a change makes the reading worse. That is the rest of this page.

A hand-graded corpus of real drawings

The test set is real isometric drawings, each one graded by a human into the bill of materials it should have produced, line by line and field by field. That graded answer is the fixture. Fixtures carry a difficulty tag and a source tag (whether the sheet is digital or scanned) so results can be segmented rather than averaged into one meaningless figure.

Fixtures come from our own sample drawings and, with written permission, from customer and design-partner packages. Nothing enters the corpus without that permission, and declining changes nothing about an engagement. The security page covers how drawings are handled generally.

Scored per field, not per drawing

A drawing is not right or wrong as a unit. Every field of every line gets scored against the graded answer for precision, recall, and F1, so a change that improves descriptions while quietly degrading schedules is visible instead of averaged away. These are the fields we track.

  • Quantity. The number that becomes a buy. A wrong quantity is a wrong order and a wrong bid.
  • Schedule. SCH 40 against SCH 40S against STD. One character changes both the price and what arrives on the truck.
  • Item number. The anchor that ties a row back to the drawing so any line stays auditable.
  • Tag. Valve and equipment tags, where present on the sheet.
  • Size. Including reducing fittings and fractional sizes, the two shapes that invite transposition.
  • Material. A106-B against A53-B, WPB against WPL6. A wrong material class.
  • Description. Scored on token-set similarity rather than exact match, because descriptions are written a dozen ways for the same part.
  • Spec code. The piping class the line belongs to, which drives the rest of the buy.

The floors a release has to clear

Each field carries a minimum F1 score. These are thresholds our own build has to clear, not predictions about your drawings. The fields that turn directly into a purchase order carry the highest floors, and description carries the lowest because it is scored on similarity rather than exact match.

Minimum F1 by field. A human reviews and approves every line regardless of what these say.
FieldMinimum F1
Quantity0.95
Schedule0.95
Item number0.95
Tag0.95
Size0.90
Material0.90
Description0.85
Spec code0.85

A change that reads worse does not ship

Every change to the extraction prompt, the model, the schema, or the PDF rendering path triggers the harness automatically. It re-reads the whole corpus and compares field by field against the last recorded baseline. A drop of more than 0.02 F1 on any tracked field fails the run, and smaller movements are reported for a human to look at. Lowering a floor is a deliberate, reviewable act with an owner's name on it, not something a change can do quietly on its way past.

Clearing these thresholds is also part of what gates a release, alongside tenant isolation testing and the rest of the ship checklist.

Measured per model, because models differ

Different AI models have different reading profiles. A run is only ever compared against a baseline recorded on the same model, because a result that beats a different model's baseline could still be a regression against its own. When we evaluate a new model, it earns its own baseline before it is allowed anywhere near production traffic.

Digital and scanned drawings are scored separately

Reading a CAD export and reading a photocopy are different problems, and reporting one number across both hides the case you actually care about on brownfield work. The harness segments results by source kind and reports the gap between them. Until the corpus holds enough of both, the report says the question is still unanswered rather than filling the gap with an average.

What specifically degrades on scanned sheets: resolution loss, skew that breaks row alignment, compression artifacts on thin strokes, and handwritten markups over the printed table. The product flags likely-scanned pages at upload so a mixed package gets triaged instead of averaged.

What the measurement does not do

It does not replace your estimator. Extraction lands as a proposal in a review screen with per-line confidence, and a person approves every line before it counts toward anything. The audit trail records who reviewed what and what they changed. That is true no matter how well a release scores.

It also does not measure something we do not do. Reading the printed bill of materials off a drawing is transcription. We do not derive quantities from drawing geometry, we do not count fittings off the picture, and we do not compute centerline lengths. Cut lengths are captured when the BOM table lists them as a row and not otherwise.

When we will publish sample results

We have not published a results table yet, and we would rather say that plainly than show you a number from a corpus too small to mean anything. When the graded corpus reaches full size, a sample table lands on this page rendered directly from the harness output that our own build is gated on, so the published figures and the figures we enforce internally cannot drift apart. It will be a record of how our corpus scored, captioned as exactly that.

The honest test is your own drawings

No published figure will tell you what your bid package does. Running a real drawing through it will. Your first project is free, no card required.

Start a project

Questions about the method? Email us. Rather have us run the first package? Managed takeoff.