Data Modelling
Delivering Semantic Architecture Through Production Engineering
Most large financial institutions face a persistent tension: the need to fix immediate engineering performance problems competes directly with the need to build long-term data architecture. The conventional wisdom treats these as separate workstreams requiring separate budgets and separate teams. This assumption is wrong.
Engineering-Embedded Data Modelling is an approach that integrates semantic modelling expertise directly into performance-focused engineering squads. Every system optimisation becomes an opportunity to implement part of a canonical data model, creating immediate operational value while building long-term strategic assets. The result is a data architecture that is proven in production from day one, not designed in isolation and adopted later.
Core Principle
A partial model, solving a business priority and embedded in production, delivers more value than a complete model sitting on a shelf.
This approach was designed and proven for a leading global investment bank facing exactly this challenge: 15 Agile teams struggling with performance and scalability, and leadership torn between funding immediate fixes and long-term data governance. The Engineering-Embedded approach resolved this tension, delivering measurable performance improvements alongside incremental FIBO-aligned canonical model coverage from Sprint 1.
For a detailed exploration of ontological data modelling and the FIBO standard, see our companion white paper: An Ontological Approach to Data Modelling.
In large financial institutions, engineering teams and data governance teams typically operate under different mandates with competing priorities. Engineering leadership needs system performance improvements now. Data governance teams need semantic consistency and regulatory alignment over time. Both are legitimate. Both are urgent. And in practice, both compete for the same budget.
Option A: Prioritise engineering performance. Quick fixes address immediate bottlenecks but create technical debt and inconsistent data structures that compound over time. Every ad-hoc fix makes future standardisation harder.
Option B: Prioritise data modelling. A canonical model is designed comprehensively, but engineering teams wait for it to be completed before seeing any benefit. Adoption risk is high: models built in isolation often fail to reflect production reality.
Both options carry significant cost. Option A accumulates hidden debt. Option B delays tangible returns and risks producing an architecture that production teams reject as impractical. The choice between them is a false dilemma.
Engineering-Embedded Data Modelling eliminates the either/or choice by integrating semantic modelling expertise directly into engineering delivery teams. Rather than treating data architecture as a separate workstream, it becomes an embedded discipline within every sprint.
Key Distinction
In this model, FIBO compliance is not an additional workstream bolted onto engineering delivery. It is a quality attribute of every engineering artefact, enforced at development time, not retrofitted later.
Each engineering pod includes a Data Modeller alongside backend developers, cloud engineers, and performance specialists. An expert Ontology Architect provides semantic guidance across pods, ensuring consistency without creating a bottleneck. This structure means that every piece of engineering work, whether optimising a query, refactoring an API, or improving a data pipeline, is simultaneously an opportunity to align that component with the canonical model.
Foundation. The Lead Ontology Architect establishes the semantic foundation by selecting relevant FIBO modules (for example: Loan, Borrower, Facility, Interest Rate, Collateral), defining organisation-specific extensions, and mapping existing system entities to the target canonical model.
Translation. Embedded Data Modellers in each pod translate ontology decisions into practical engineering artefacts: database schemas, API specifications, domain objects, validation rules, and code templates that enforce semantic constraints. Engineering teams do not need to understand ontology theory. They receive clear, usable patterns.
Continuous Alignment. Every sprint includes model alignment verification. The Ontology Architect reviews planned work for alignment opportunities. Embedded specialists ensure new code follows canonical patterns. Sprint reviews report model coverage metrics alongside performance improvements.
The recommended delivery unit is high-performance Agile pods, each comprising specialists, supported by a shared Ontology Architect.
| Role | Contribution |
|---|---|
| Backend Developer | Java/Spring Boot specialist delivering feature and performance work against FIBO-compliant patterns |
| Cloud/Database Engineer | AWS, Snowflake, or equivalent platform optimisation with semantically aligned schema design |
| Data Modeller / DQ Specialist | Translates ontology decisions into schemas, validation rules, and code templates; enforces data quality |
| Performance Engineer | System optimisation with end-to-end visibility through model-based lineage tracking |
| Scrum Master / Technical Lead | Delivery coordination ensuring both performance and model objectives are met each sprint |
Pods are supported by an Expert Ontology Architect who provides semantic governance, reviews alignment across workstreams, and ensures consistency as model coverage expands.
Prospect 33 delivers these teams through its nearshore development centres, combining significant cost efficiency with access to a rapidly growing fintech talent pool. Time zone overlap with both European oversight and global banking operations enables real-time collaboration without the coordination overhead of distant offshore models.
Outcome
Engineering performance improvements, FIBO-compliant canonical model coverage, and reduced technical debt — delivered through a single integrated engagement rather than competing workstreams.
Engineering teams working with FIBO-aligned schemas gain clear data contracts with unambiguous entity definitions, reducing time spent deciphering inconsistent structures. Consistent canonical patterns eliminate custom mapping logic between systems. Built-in data quality checks catch issues at development time rather than in production. Reusable semantic components accelerate new feature delivery and reduce rework.
The canonical model is built where data actually flows, not in a theoretical design document. This means the model is validated through real usage from its first implementation. Adoption is incremental rather than big-bang, eliminating the risk of a completed model that production teams reject. FIBO alignment ensures readiness for regulatory reporting requirements, and the model can be extended with the FIBO Regulatory Reporting Ontology (FRO) as compliance needs evolve.
Machine-readable semantics built into production data from day one provide a foundation for intelligent data processing. Clean, semantically consistent structures reduce the data preparation burden that typically consumes the majority of analytics effort. The FIBO ontology enables AI systems to reason with context rather than raw data points, supporting advanced capabilities from anomaly detection to automated regulatory validation.
Model-based lineage tracking provides end-to-end visibility into data processing chains, enabling root cause analysis that accounts for upstream and downstream dependencies. Semantic understanding ensures that optimisations in one system do not create problems in connected systems. Performance improvements benefit multiple consumers of the same canonical data structures.
Any approach that integrates two disciplines into a single delivery model introduces specific challenges. Acknowledging these transparently is essential to effective implementation.
Embedding data modelling into engineering sprints will reduce initial velocity as teams learn FIBO-compliant patterns. This is a deliberate investment. The structured training programme (see Appendix A) mitigates this through progressive skill building: teams gain ontology fundamentals in the first two weeks, apply practical patterns in weeks three and four, and reach independent proficiency within two months. Velocity typically recovers and exceeds baseline within the first quarter as reusable patterns reduce rework.
FIBO and ontology expertise is a niche skill set globally. The delivery model addresses this by concentrating specialist knowledge in the Ontology Architect role, which provides guidance across both pods, while embedded Data Modellers translate that guidance into practical engineering artefacts. Engineering team members do not need to become ontologists; they need to follow well-defined patterns. However, the Ontology Architect role is a single point of dependency and should be supported by documented decision frameworks and knowledge transfer protocols.
Introducing new pods operating under a different methodology alongside established Agile teams creates potential friction. Clear governance protocols, shared sprint ceremonies where appropriate, and explicit interface agreements between new and existing teams are essential. The pods should be positioned as complementary capability, not replacement.
FIBO model coverage grows organically based on engineering priorities, not predetermined percentages. This is a feature, not a limitation: it ensures that every modelled component is production-tested and practically validated. However, stakeholders should understand that coverage is determined by which systems are prioritised for engineering work, not by an abstract modelling roadmap.
This approach was designed for a leading global investment bank with 15 Agile teams facing significant performance and scalability challenges across critical loan processing systems. Leadership was under pressure to demonstrate both immediate engineering improvements and progress toward a FIBO-aligned canonical data model for their Loans Enablement Platform.
The Engineering-Embedded solution resolved what had previously been framed as a binary budget allocation decision. By integrating data modelling expertise directly into performance-focused pods, the engagement could deliver measurable system improvements from the first sprint while building canonical model coverage incrementally across every area touched by engineering work. The model is not adopted after completion; it is being built in production from day one.
Organisations that pursue engineering performance and data modelling as separate engagements typically face compounding costs: separate teams, separate governance, separate stakeholder management, and the integration overhead of reconciling independently developed outputs. The Engineering-Embedded approach consolidates these into a single delivery stream.
Combined with Prospect 33's nearshore delivery model, the integrated approach typically achieves significant cost reductions compared to running parallel workstreams — in many cases effectively delivering the canonical data model at no incremental cost above the engineering engagement. Detailed commercial proposals are tailored to each client's specific requirements and scale.
Time to value is similarly accelerated. Traditional sequential approaches require 18–27 months before the data model delivers production value. The Engineering-Embedded approach produces value from month one, with capability realised every sprint thereafter.
Prospect 33 has deep capability in the intersection of data modelling, engineering delivery, and financial services domain expertise.
Access to doctorate-level ontologists with FIBO implementation experience across major banking institutions.
Highly competitive rates for experienced data modellers and data quality specialists, delivered through nearshore centres of excellence.
Proven Agile pod formations with Java/Spring Boot, cloud platform (AWS, Azure), and database optimisation expertise.
Domain experience spanning loans, trade surveillance, regulatory reporting, risk management, and operational processes across Tier 1 and Tier 2 institutions.
Significant depth of experienced data scientists, enabling seamless progression from semantic data foundations to AI-driven capabilities.
The choice between funding immediate engineering needs and investing in long-term data architecture is a false dichotomy. Engineering-Embedded Data Modelling delivers both through a single, integrated approach: every system optimisation simultaneously advances the canonical model, and every model implementation is validated in production from day one.
The result is faster engineering delivery, lower technical debt, production-proven data architecture, and a semantic foundation that enables AI, analytics, and regulatory compliance. Organisations that adopt this approach do not choose between velocity and architecture. They achieve both.
Embedding FIBO-aligned practices into engineering teams requires structured skill development. The following programme is designed to build practical proficiency without disrupting delivery momentum.
The following example demonstrates how FIBO-aligned engineering produces cleaner, more maintainable code compared to traditional ad-hoc approaches. This is illustrative of the patterns provided to engineering teams through the code template library.
Inconsistent, undocumented data structures that vary across systems and require manual interpretation:
// Ad-hoc loan entity — inconsistent across systems
public class LoanRecord {
String loanId;
String borrowerName;
Double amount;
String productType; // Inconsistent values across systems
Date startDate;
}
Self-documenting, type-safe structures with built-in validation and cross-system consistency:
// FIBO-compliant loan entity with semantic validation
@FIBOEntity(concept = "https://spec.edmcouncil.org/fibo/ontology/LOAN/")
public class Loan extends FIBOFinancialInstrument {
@FIBOProperty(property = "hasLoanIdentifier")
private LoanIdentifier loanId;
@FIBOProperty(property = "hasObligor")
private BorrowerParty borrower;
@FIBOProperty(property = "hasPrincipalAmount")
@DataQualityRule(rule = "POSITIVE_AMOUNT")
private MonetaryAmount principalAmount;
@FIBOProperty(property = "isClassifiedBy")
private LoanProductType productType; // Controlled vocabulary
@FIBOProperty(property = "hasEffectiveDate")
private BusinessDate originationDate;
}
Type Safety: FIBO-compliant types prevent data type errors at compile time, eliminating a class of production defects.
Built-in Validation: Data quality rules (e.g., POSITIVE_AMOUNT) catch issues at ingestion rather than downstream processing.
Cross-System Consistency: The same loan representation is used across all systems, eliminating custom mapping logic.
Interoperability: Standard FIBO properties enable integration without point-to-point translation layers.
Self-Documentation: Semantic annotations make code self-documenting, reducing onboarding time for new team members.
The following terms are used throughout this document. Definitions are provided for non-technical stakeholders.
| Term | Definition |
|---|---|
| Canonical Model | A single, authoritative representation of data entities and their relationships, used as the standard across all systems in an organisation. |
| FIBO | Financial Industry Business Ontology. An open standard that defines financial concepts with semantic precision, maintained by the EDM Council. |
| FIB-DM | FIBO-derived data model. A practical transformation of FIBO into relational database structures that engineering teams can implement directly. |
| FRO | FIBO Regulatory Reporting Ontology. An extension of FIBO that maps data model concepts to regulatory reporting requirements. |
| Ontology | A formal definition of concepts, their properties, and relationships within a domain. Goes beyond structural modelling to capture meaning. |
| Semantic Layer | A translation layer that maps business terminology to underlying data structures, enabling consistent interpretation across users and systems. |
| Data Lineage | The ability to trace data from its origin through every transformation and system it passes through to its final use. |
| Technical Debt | The accumulated cost of quick-fix engineering decisions that must eventually be corrected, analogous to financial debt that accrues interest. |
| Agile Pod | A small, cross-functional team (typically 5–7 people) that works in short iterative cycles (sprints) to deliver working software incrementally. |
| Sprint | A fixed-length development cycle, typically two weeks, at the end of which the team delivers tested, working functionality. |
For more information on how Engineering-Embedded Data Modelling can transform your organisation's approach to data architecture and engineering delivery.
www.prospect33.com