Every metal fabrication project starts the same way: a folder of PDFs. General arrangement drawings, isometrics, detail sheets: vector exports from the design office, or scans of markups that came back from the field. Somewhere inside each one is a bill of materials. Getting that table into a spreadsheet you can order from is called takeoff, and on most projects it is still done by hand.
What "BOM extraction" means
A bill of materials extracted from a drawing isn't just "the numbers on the page." A usable BOM line has to preserve five things:
What the item is: a beam profile, a pipe fitting, a plate.
Grade and standard. A wrong substitution is a certification failure, not a typo.
Length, size, schedule, thickness: whatever the item's geometry requires.
And the unit it's counted in: each, metres, linear feet.
Which drawing and which item number the line came from, so a question six weeks later ("where did this come from?") has an answer.
Drop any of these and the BOM stops being usable for procurement. You cannot order against a line that says pipe, qty 13 with no size and no schedule.
Why PDFs make this hard
A PDF is a rendering format, not a data format. Three properties of real-world drawing sets make extraction difficult rather than a trivial parsing problem.
A drawing exported straight from CAD has selectable text, so the table can, in principle, be read directly. A scan, or a PDF of a markup that went through a printer at some point in its life, has none: the "text" is pixels. Reading it means OCR, with everything that implies about low-resolution numerals and near-identical characters.
There is no standard for how a title block or a material take-off table is laid out. Every drafter, every project, sometimes every revision, arranges columns differently, uses different abbreviations, orders items differently. A tool, or a person, has to work out which column means quantity and which means size from context, not from a fixed template.
The same drawing number exists at rev A, rev B, rev C, with real changes buried in a table that otherwise looks identical. Takeoff against a superseded revision doesn't fail loudly. It produces an order for the wrong quantity of the wrong thing, and nobody notices until the material shows up.
| DESC | SEAMLESS PIPE 4" |
| SCH | 40 |
| MATL | A106-B |
| LGTH | 6.40 m |
| ITEM | Steel pipe 4 inch |
| SPEC | SCH40 |
| GRADE | A106 Gr.B |
| QTY | 2.15 m |
| NO. | TUBE S/S DN100 |
| — | SCH 4O |
| — | A1O6B |
| QT | 3.80 m |
The actual cost of doing it by hand
Manual takeoff is a real skill, and experienced estimators are fast at it. But the cost isn't the time spent reading. It's what happens when volume goes up. A project with three drawings gets read carefully. A project with sixty drawings, on a deadline, gets read tired, at the end of a long day, by whoever is available. That is where transcription errors live: a quantity off by one digit, a schedule read as 40 instead of 80, a line skipped because it sat on page 2 of a drawing everyone assumed was single-page.
Those errors don't surface at takeoff. They surface at delivery, when the wrong material is on the truck.
What good automated extraction looks like
The bar isn't "faster than a human." It's producing a BOM an estimator can review faster than they could have built it from scratch. That means every line stays traceable back to its source drawing and item number, so spot-checking against the original PDF takes seconds rather than a re-read of the whole set. Confidence should be visible per line, not hidden: a line read from clean vector text and a line read from a blurry scan carry different risk, and hiding that distinction just pushes the error further downstream.
- A flat spreadsheet with no source column
- One global accuracy figure
- Silent auto-correction of odd values
- Drawing + item number on every line
- Per-line confidence, visible
- Ambiguities flagged for review, not fixed quietly
Where extraction fits in the bigger picture
Extraction is the first step, not the whole job. A raw extracted BOM from a real drawing set will have the same item worded three different ways across three drawings, and it needs a cutting plan before anyone can order material against it efficiently. That's a separate problem, covered in BOM consolidation: why the same part gets ordered three times, but one that depends entirely on extraction producing clean, traceable lines to begin with.
Every item stays traceable to its source drawing and item number, then carries into consolidation and cutting optimisation.
Frequently asked
Yes, through OCR, but the risk profile differs from vector text. Character confusions on numerals are the common failure, so scanned lines should be flagged with lower confidence and reviewed against the source sheet.
Takeoff is the whole estimating activity: reading drawings and quantifying material. Extraction is the mechanical part of it: getting the tables that already exist on the sheets into structured, orderable data.
By keying lines to drawing number and revision, so a superseded sheet in the same folder is visible as a conflict rather than merged silently into the totals.