By Alkis Papadopoullos, CEO and CTO of Coginov
You can get to a working invoice extractor in a quarter. One engineer, an LLM API, a JSON schema, a validation pass on top, and a demo that makes the room nod. If your document types are narrow, your layouts are homogeneous, and every invoice, PO and BOL your customers ever touch flows into exactly one schema you control — go build it. The math works. The demo will land. I’d probably build it too.
That’s the honest concession, and I want to get it out of the way before the rest of this argument, because the rest of this argument is why almost none of the ERP and CRM vendors currently having the build-vs-buy conversation actually live inside those constraints — and why the thing that kills these projects is never the extractor.
The clock-speed problem first, because it’s the one nobody costs into the business case. The model layer moves every few months. Not the vendors’ marketing cycles — the underlying capability. A frontier model release in Q2 quietly resets what “good” means on semi-structured extraction, and the delta compounds. Your ERP release train, meanwhile, moves every 12 to 24 months, with certification gates, regulatory review in regulated verticals, and a customer base that upgrades on its own schedule regardless of what you shipped. When you build the extractor in-house, you have just coupled a fast-moving component to a slow-moving one. The extractor’s ceiling is now the release train’s cadence. Two years in, your extraction quality is frozen relative to the frontier, and — this is the part that hurts — nobody on your team is measuring the gap, because the thing still passes the same acceptance tests it passed on day one.
Which brings me to what actually costs money, and it isn’t the extractor. It’s the apparatus around it. Four categories, and I’m going to name them precisely because naming them is the point of this post — a team evaluating build-vs-buy needs to know what they’re actually signing up to own.
Confidence calibration. The model returns a value for “invoice total.” Do you trust it enough to auto-post, or does it go to a human queue? That decision cannot be “the model said so” and it cannot be a fixed threshold. It has to be calibrated per field, per document class, per tenant, and it has to be re-calibrated every time the underlying model changes. Getting this wrong in the auto-post direction creates silent financial errors in your customers’ books. Getting it wrong in the queue direction floods your customers’ AP clerks and they turn the feature off. There is no comfortable middle you find once and leave alone.
Exception routing. When the model is not confident, or when validation against the ERP’s own master data fails (vendor not on file, PO number doesn’t match, tax code inconsistent with jurisdiction), the document has to go somewhere useful. Not a generic “review queue” — a routed queue with the right context attached, to the right role, with the fields flagged that actually need human attention. This is a workflow product hiding inside your extraction product, and it is where user love or user abandonment is decided.
Per-tenant layout drift. Your customer’s top-20 vendors change their invoice templates on their own schedule. A logistics provider redesigns their BOL in March. A supplier switches ERPs in June and the PO acknowledgments now come out of a different system with different field placements. Extraction accuracy on that tenant degrades the following week, and if you don’t have per-tenant monitoring you learn about it from a support ticket three months later — after several thousand invoices have been mis-posted. Fleet-wide accuracy metrics hide this. You need the per-tenant view or you’re flying blind.
A regression harness. This is the one that scares me most in build scenarios, because it’s the one teams skip. When a new model version comes out — yours, the vendor’s, doesn’t matter — and someone wants to upgrade, you need to be able to run the candidate against a labelled, versioned, tenant-representative test set and see whether accuracy went up, went down, or moved sideways with a shift in error distribution that changes which fields fail. Without this harness, “upgrade the model” is a coin flip that silently degrades production. With it, model upgrades become a routine engineering event. This apparatus is not glamorous. It is also not optional.
We’ve built each of these inside QoreCapture across invoices, POs and BOLs, with per-tenant tuning and validation queues, and I will tell you plainly that the harness discipline took longer to get right than the extraction pipeline itself. That ratio — apparatus-to-extractor — is the number that build-scenarios systematically underestimate.
Second asymmetry, and then I’ll close. Your customers’ documents don’t only flow into your ERP. Their AP team receives invoices that also need to hit their expense system, their tax provisioning tool, their bank reconciliation. Your CRM’s contract intake feeds legal review, procurement, and revenue recognition. The extractor that only knows how to talk to your schema is a feature; the extractor that participates in the customer’s actual document estate is a platform. Your resellers already know this, because they sell into mixed estates and they hear the objection in every deal. This is the part of build-vs-buy that gets waved away with “we’ll add integrations later” — and later, in practice, means an MCP-style connector layer that becomes its own product line, competing with your roadmap for the same engineers.
So the practical takeaway, for someone reading this on Monday morning with a build-vs-buy memo open in a tab.
Cost the apparatus, not the extractor. If your business case only prices the LLM calls and the prompt work, you have priced roughly 20% of what you’re actually going to own. Ask your team, specifically, who owns the calibration recalibration cadence, who owns the per-tenant regression detection, who owns the harness that lets you upgrade the model without a war-room, and who owns the connector surface that lets the same extraction serve documents flowing out of your product. If the answers are “we’ll figure that out post-launch,” you don’t have a build plan, you have a demo plan. Those are very different artifacts, and the market decides which one you shipped somewhere between month six and month eighteen — usually quietly, in the churn cohort.
Visit our new QoreUltima Solutions’ page.
We create innovative solutions
COGINOV is recognized as a world leader in semantic technologies and information management. We are a Canadian software company offering our customers innovative solutions for managing structured and unstructured information. Our head office is based in Montreal.
Coginov’s Qore platform technology enhances the information value chain, transforming unstructured content into highly contextualized, accessible and valuable information. Coginov’s solutions enable you to capture, analyze, engage, automate and manage your information assets, with unrivalled accuracy and efficiency.
Discover our solutions QoreAudit, QoreUltima and QoreMail
2022 Marketing. All Rights Reserved by Artureanec