Ark, One Year On: The Architecture Held

What a year of building taught us about prevention — and the identity underneath it

October 5, 2026

Eleven months ago I published an argument, and it traveled further than anything I’ve written.

The argument was simple. Every legacy master data governance tool — SAP MDG, Oracle MDM, the modern alternatives — shares one architectural assumption: that governance happens after data entry. They put guardrails around dirty data that’s already in your ERP. I compared it to building a water treatment plant downstream while the factory upstream keeps dumping toxins. You treat symptoms. The pollution never stops.

So we built the other thing. Governance at the point of creation. Prevention, not remediation.

(That piece — Why We Built Ark — is still the most-read thing we’ve published. Which tells me the argument landed. This is the sequel to it.)

Making the argument was the easy part. Then we spent a year building it out. Here’s what building it actually taught me — and what Ark does now that I couldn’t have promised you last year.

The thesis held — but it was bigger than I said

When we started, “prevention” mostly meant one thing: stop a technician from creating a duplicate part at the moment of request. That still works. It’s still true.

But building it taught me prevention isn’t a feature you attach to parts data. It’s a discipline that runs across the entire MRO master — because bad data doesn’t respect the boundaries between your item catalog, your vendor list, and your asset register. A duplicate vendor corrupts spend analytics exactly the way a duplicate part corrupts inventory.

And underneath all of it was one thing I hadn’t named in the first piece. Everything prevention needs — stopping the duplicate, enriching the new part, cleaning the old ones — depends on the system knowing what a thing actually is. MRO data almost never does. It has descriptions. It has the words a hurried technician typed. It has never had an identity.

So that’s what the year came down to: an AI engine, grounded in the MRO domain, that gives every part the identity it never had — a platform that puts that identity to work across every master, with the provenance to prove every call it makes.

One — The identity, built at creation

Ask a technician to create a new part and you’ll get a terse, inconsistent description typed under pressure while a machine is down. That’s the largest single source of dirty MRO data. Not carelessness — friction.

Point that request at a manufacturer part number and Ark does more than look up specifications. It researches and fills the technical attributes, each carrying its source — and then it classifies the part: not by the keywords in a description, but by what the part actually is, against Ark’s MRO taxonomy. From there it guides the operator through the fields that remain, in the order the part’s own class demands, until the record is complete and correct.

The technician never becomes a data-entry clerk. They’re led, not left. And the part is born with something MRO data almost never has at the moment it’s created: an identity — a classified, attribute-complete record, not a line of text typed in a hurry.

One thing mattered enormously in building this, and it’s worth saying plainly. We do not let the AI invent your master data. A generated specification that’s right most of the time is worthless for a part you’ll purchase against for a decade. So Ark’s engine is grounded — in the MRO domain, in your taxonomy, in proven methods that return a fact with a reason attached, not a confident guess. Generation proposes. Grounded intelligence and the operator decide.

That is the difference between an AI you can put in front of a master data record and one you can’t.

Two — The duplicate that never happens

Everyone in this industry has run the quarterly de-duplication project. By the time you find a duplicate, it has already done its damage — ordered against, inflating inventory, skewing a report someone already decided on. Finding it late doesn’t undo any of that. It deletes the evidence.

The only duplicate that costs nothing is the one that never gets created.

Here’s why that was impossible before — and why identity changes it. For as long as MRO has existed, “finding a duplicate” meant matching descriptions. Keyword or fuzzy, it’s the same trick underneath: measure how much one line of text reads like another. But the description is the one field you cannot trust. The same part is SCREW in one record and SCRW in the next. Sometimes it has no description at all. Match the text and you miss the twin standing right beside it.

Ark doesn’t match the text. Because every part now has an identity — what it is, not how it was spelled — Ark compares on the facts. Add-on equals add-on. 690V equals 690V. Screw terminal equals screw terminal. The label is noise; the identity is exact. It catches the duplicate hiding under a different abbreviation, a different supplier, a different part number — even one with no description to match against at all.

Understand the part first. Compare second. And do it before the record is committed — configurable by policy to warn the operator or hard-block the creation outright, with a logged supervisor override when human judgment is warranted. Not a report you act on next quarter. A duplicate that never enters. And because bad data ignores the walls between your masters, Ark does this across item, vendor, and asset alike.

(I went deep on exactly how this plays out — the SCREW/SCRW catch, the eight identical parts with blank descriptions — in The Identity MRO Never Had. This is the short version.)

Three — The mess you already have

Here’s the part most AI tools skip, and the part that matters most for master data.

An AI that hands you a confident answer with no way to check it is a liability in a system you’ll purchase, maintain, and make safety decisions against for a decade. “The model said so” is not governance.

So every value Ark produces carries its provenance — where it came from and how confident Ark is in it — and Ark is honest about the difference between a specification it researched and verified against a source and one it reasoned to on its own. Nothing arrives as an anonymous fact you’re asked to trust blind.

The same holds for the verdicts. When Ark flags a duplicate, it doesn’t give you a similarity score and a shrug. It gives you the reason: this fact equals that fact, and here is the record that proves it. A fuzzy match tells you it’s 87% sure. Provenance tells you why — in a form that stands up to an audit, whether you’re blocking a creation or letting a supervisor wave one through.

That is the line between AI you can put in front of a master data record and AI you can’t. The intelligence is what makes Ark fast. The provenance is what makes it trustworthy. You need both — and almost nothing else in this category has the second one.

One platform, three masters — and what only appears when they're all clean

A year ago, Ark really spoke to one thing: parts. Today all three MRO masters are built, complete, and running end-to-end — Item, Vendor, and Asset — and, more to the point, linked. Items connect to the vendors that supply them and the equipment they serve. Equipment connects to its own failure and problem history.

Three clean masters is worth doing on its own. But the reason to build all three isn’t three clean datasets. It’s what appears only once they’re clean and connected — the questions you simply could not ask before, because the data to answer them sat in three silos, each too dirty to trust.

Now Ark answers them directly:

  • Where am I exposed? The parts only a single supplier makes, sitting under the equipment you can least afford to have down.
  • What actually fails — and what does it take to fix? The most common failure patterns for each equipment type, with the spare parts and vendors behind them.
  • What does this machine really need? For any equipment type, one view of the parts it consumes, who supplies them, and what can go wrong.

None of that is possible on a dirty foundation. All of it becomes routine once identity runs across the whole master. That was the real icing on the year — the payoff for treating prevention as a fabric, not a feature.

(What “good” actually means for each, I set out master by master earlier this year — item, vendor, and asset.)

The quieter lesson: Ark reads your data before it governs it

None of this works if Ark assumes what your data looks like — because no two organizations’ MRO masters are shaped the same. Some are rich with structured attributes. Some are barely a short description and a prayer. So before Ark establishes identity, it orients itself to the actual shape of your data and adapts. No configuration screen where you declare your structure. No assumptions that shatter on contact with a real catalog. Ark comprehends first, then acts.

This was the deepest lesson of the year, and the least visible. Generic tools can’t do it — they were built to assume one shape, and MRO doesn’t have one.

Deployable on your terms

Last year Ark was “proven with leading manufacturers.” A year on, it’s deployable however your enterprise needs it — including fully air-gapped, on-premise environments for organizations that can’t send MRO data to the cloud, and tiered so you can start where the pain is sharpest and expand across masters from there.

The proof kept accumulating. 50,000+ duplicate parts eliminated before a single migration. 270,000+ master parts delivered with complete specifications. Sixty-day implementations against the twelve-month legacy standard. Not slideware — deployed, in production, holding.

The thing I still believe — more, not less

In the first piece I wrote that data governance isn’t a technology problem. It’s an engineering problem. A year of building only made me more certain.

Generic MDG stalls on MRO data because it was built for simpler domains — product catalogs, customer records — that assume a shape and a rhythm MRO doesn’t have. MRO has technical specs that matter, manufacturer part numbers that are non-negotiable, classification that has to serve maintenance and not just procurement, and structure that varies wildly from one plant to the next. You cannot govern that with assumptions. You engineer for it — you give the data an identity it never had, and the provenance to prove it — and then you stay to make it work.

That’s still the difference. We don’t hand you a report. We build the solution, and we stay.

Where this goes

The industry is at the same inflection point it was a year ago — S/4HANA migrations, post-M&A consolidation, AI and predictive maintenance that all die on dirty data. None of that has changed.

What’s changed is the answer. It’s a year more built, a year more proven, and a great deal broader than the parts catalog we started with. The argument was that prevention beats remediation. Building it didn’t just confirm that. It showed me what prevention actually required underneath: an identity for data that never had one — and the provenance to trust it.

Clean data was never meant to be a project you run every quarter. It’s meant to be the only kind of data your system can create — because every part carries an identity, and a provenance, from the moment it exists. That’s the architecture. A year on, it held.

The year, in the writing

I didn’t announce these as features. I wrote through each one as I built it. If you want the long version of any part of this piece, the trail is here.

The argument

Why generic AI can’t do this

Identity, and the duplicate that never happens

Provenance — AI you can actually trust

What “good” looks like, master by master

Reading the data before governing it

Why we build it to stand alone — and why we stay

 

See what Ark reveals about your data

Start with a complimentary data quality assessment. We’ll analyze a sample of your MRO data and show you what’s hiding in it — duplicates, missing attributes, classification gaps — and what it’s costing you annually.

No obligation. No sales pressure. Just a clear picture.

→ Get Free Data Assessment: sales@bluemindz.com

→ Request Ark Demo: sales@bluemindz.com

About the Author

Raghu Vishwanath

Raghu Vishwanath has spent thirty years in the trenches of enterprise data — across financial services, manufacturing, utilities, retail, high-tech, and pharma — turning messy business reality into systems software can actually act on. He is Managing Partner at Bluemind Solutions, a product engineering firm specializing in MRO master data governance, and writes about software engineering, AI, and building platforms that last.