How your documents become a foundation.
The same method whether you come from construction, building services or industry. An automated pipeline does the building, your experts decide. No workshop marathons, no modelling by hand.
Five steps, your people in the loop at every one.
The video walks through the story end to end. Below is the machinery inside each step. One thing holds throughout: an automated pipeline does the building, and your experts appear exactly three times, at the approval gates, to check what it proposes rather than model anything themselves.
- 01
Start
Ontology design starts from your industry's open standard: for finance, for construction, for manufacturing. The pipeline loads and maps it automatically: entity types, properties and permitted relationships, thousands of terms your industry has already agreed on. That draft becomes your ontology, the formal vocabulary the whole graph is built on. Nobody on your side models anything; your experts will only correct it later.
Under the hood- The standard is loaded as the ontology skeleton (TBox): classes, properties and the relationships permitted between them.
- The standard is the prior, not your documents: it carries distinctions regulation requires that your data never wrote down.
- It loads from a local bSDD/OWL export or the bSDD API, mapped straight into the graph schema. No modelling by hand.
DeliversA draft ontology for your sector, anchored in a published standard, before a single internal meeting. - 02
Surface
Nobody fills in templates and nobody sits in workshop marathons. The pipeline scans your documents and systems and proposes two things. First the competency questions: the concrete questions the graph must be able to answer, like “which contracts renew next quarter, and on what terms?”, each stored with its expected answer as a permanent test case. Second the definitions, exceptions and rules it finds behind them, each with its source attached. Your experts review a finished list in one sitting. This is where knowledge graph projects usually stall for months; here it is the fast part, because the drafting is machine work.
Under the hood- Each competency question is stored with its expected answer and a test query the finished graph has to satisfy.
- Two elicitation paths, one output: a guided expert interview and mining your documents both produce the same question set.
- Definitions and exceptions are surfaced from your data, each carrying the source document and page it came from.
DeliversA set of competency questions with expected answers, plus every definition and exception: proposed by the pipeline, reviewed by your experts. - 03
Confirm
Before any data flows, the draft ontology is tested against every competency question: can it answer each one, in the right shape? Then comes the human gate, and this is the whole of what we ask of your experts: review what the pipeline proposed, correct it where the business disagrees, approve. Nothing enters the until a named expert has signed it off. This gate is deliberate: a wrong definition that looks official does more damage than a missing one. Every element of the ontology exists to answer a real question, and every fact traces back to the person who approved it.
Under the hood- Ontology evaluation: every competency question runs against the schema. Can it answer each one, in the right shape?
- SHACL rules act as a deterministic gate: anything malformed is caught by machine, so a human only ever sees valid candidates.
- The human gate is one of exactly three: a named expert reviews, corrects and signs off before anything counts as true.
- Each decision is stored with its owner and timestamp, so the authoritative definition can always be traced to a person.
DeliversAn evaluated, expert-approved ontology. Every element serves a question; every fact has a named owner. - 04
Connect
Now the pipeline fills the graph, unattended. Connectors read your documents and systems (PDF, Word, Excel, SAP exports, IFC models), and extraction runs in two passes: first the , then the relationships between them, stored as triples. Every triple carries its source document, page and timestamp. Duplicates are resolved: the same customer under three names in three systems becomes one record with a stable identifier. Nobody on your side keys anything in; at the second gate your experts review a sample of the extracted facts, one by one, before the full load runs.
Under the hood- A connector per format: PDF, Word, PowerPoint, Excel, email, IFC/BIM models, SAP exports. Scanned pages go through OCR.
- Two-pass extraction: entities first, then the relationships between them, written as triples (subject, relation, object).
- Entity resolution: three names in three systems collapse to one stable identifier, so your AI stops counting them as three.
- Provenance on every triple: source document, page, timestamp and a confidence score. This is what makes each answer traceable.
- A sample-validation gate before the full load: your experts review a batch of extracted facts, one by one.
DeliversA populated knowledge graph: every fact traceable to document and page, every real-world thing represented exactly once. - 05
Keep true
A graph no one maintains drifts away from the business it describes. So maintenance is a pipeline, not a promise. When a source document changes, it deletes exactly the facts derived from it and re-extracts that one document, nothing else; the provenance on every triple is what makes that targeted update possible. After every change to the ontology, all competency questions run again as regression tests, automatically. And every version is kept, so you can see what changed, and when.
Under the hood- Delta re-ingestion: change one document and only the facts derived from it are deleted and re-extracted. The provenance makes this surgical.
- CQ regression: after any change to the ontology, every stored competency question runs again, automatically, as a test.
- Versioning: the ontology carries a version tag and every version is kept, so you can see what changed and when.
- Temporal confidence: facts carry a valid-from and valid-to, and trust decays over time, so stale meaning is flagged, not trusted.
DeliversA layer that stays correct as your business changes: targeted updates, automatic re-testing, full version history.
See it on a question from your own company.
The five steps above never change. Where they start for you is what a first conversation is for.
