Ask a mid-sized NBFC how many AI models it has in production and the answer is frequently uncertain. Models arrive through vendor products, embedded in platforms, built by individual teams. Nobody holds the complete list. That gap — not model accuracy — is where most institutional AI risk actually sits.
What Model Risk Is
Model risk is the risk of adverse outcomes from decisions based on models that are incorrect, misused, or applied outside their intended purpose.
It arises three ways, and institutions usually focus on only the first:
| Source | Example |
|---|---|
| The model is wrong | Poor design, unrepresentative data, flawed assumptions |
| The model is used wrongly | A model built for one segment applied to another |
| The model has stopped being right | Drift since deployment |
The second is underappreciated. A model validated properly for salaried applicants and then used on self-employed applicants has not failed — it has been misapplied. No amount of technical validation prevents this; only governance does.
The Model Inventory
Everything else depends on this, and it is where most institutions start behind.
A complete inventory records every model in use, including ones the institution did not build. For each:
| Field | Why it matters |
|---|---|
| Purpose and permitted use | Defines what counts as misapplication |
| Owner | Named accountability, not a team |
| Build or vendor | Determines validation approach |
| Data sources | Enables impact assessment when a source changes |
| Risk tier | Governs validation and monitoring intensity |
| Validation date and outcome | Evidence for audit |
| Monitoring in place | Whether drift would be detected |
| Last review | Identifies models nobody has looked at |
The practical difficulty is shadow models — spreadsheets with embedded logic, vendor features enabled without procurement review, scripts built by a team and quietly relied upon. These make decisions and appear in no register.
Risk Tiering
Not every model warrants the same scrutiny. Tiering allocates effort sensibly.
| Tier | Characteristics | Treatment |
|---|---|---|
| High | Affects customer outcomes directly — credit decisions, pricing, fraud blocks | Independent validation, continuous monitoring, board visibility |
| Medium | Influences decisions with human review — collections priority, marketing | Validation, periodic review |
| Low | Internal efficiency with limited customer impact | Proportionate review |
The tiering criterion should be customer impact, not technical sophistication. A simple rule-based scorecard declining loan applications is higher risk than a complex model optimising internal document routing.
Validation
Validation must be independent of the team that built the model. A developer validating their own work is reviewing, not validating.
What it covers:
- Conceptual soundness — is this approach appropriate for this problem?
- Data quality and representativeness — including whether the training population matches the application population
- Performance — on data never used in development
- Stability — behaviour across segments and conditions
- Fairness — outcome disparities across groups
- Explainability — can decisions be accounted for
- Limitations — documented conditions under which the model should not be used
That final item is the most neglected and among the most useful. A validation report stating “not validated for applicants below a given income band” gives the business a clear boundary. Without it, the model gets applied wherever someone thinks it might work.
Third-Party Models — the Hardest Part
Most institutions buy more models than they build. Under RBI’s FREE-AI framework, accountability stays with the regulated entity regardless of who built the model.
This creates a practical tension. Vendors resist disclosing methodology as proprietary. Institutions are nonetheless expected to validate.
| Approach | What it achieves |
|---|---|
| Contractual disclosure rights | Access to methodology, data sources and testing |
| Outcome testing | Validate performance and fairness even without internal visibility |
| Challenger comparison | Benchmark the vendor model against an internal alternative |
| Change notification | Requirement to inform you when the model is updated |
| Right to audit | Escalation path where concerns arise |
The fourth row is routinely missing from contracts and matters considerably. A vendor updating its model without notification means the institution is running an unvalidated model while believing otherwise.
Outcome testing is the practical fallback: even where methodology is opaque, you can measure what the model does across segments. That reveals fairness problems and performance decay without needing to see inside.
Monitoring
Validation at deployment is a point-in-time exercise. Models degrade, so monitoring is what keeps validation meaningful.
| Monitor | Detects | Speed |
|---|---|---|
| Input distribution shift | Data drift | Immediate |
| Output distribution shift | Changing behaviour | Immediate |
| Leading indicators | Early performance decay | Weeks to months |
| Outcome performance | Definitive degradation | Slow |
| Fairness metrics | Emerging disparity | Ongoing |
| Override rates | Staff losing confidence in the model | Immediate |
The last row is an underused signal. When staff increasingly override model recommendations, they have observed something the metrics have not yet surfaced.
Governance
RBI’s framework points toward board-level ownership rather than treating AI as a technology matter.
- Board-approved AI policy defining appetite and prohibited uses
- Named accountability per model — a person, not a function
- Model risk committee reviewing high-tier models and approving deployment
- Internal audit coverage of the AI estate
- Documented escalation when monitoring triggers fire
- Capacity building so those overseeing models can meaningfully do so
That last point is a genuine constraint. Boards and audit functions are being asked to oversee models few of their members can evaluate technically. Building that capability is slower than buying the models.
Key Takeaways
- Model risk arises from models being wrong, misused, or no longer right
- Misapplication is underappreciated — governance prevents it, validation does not
- Everything depends on a complete model inventory, including shadow models
- Tier by customer impact, not technical sophistication
- Validation must be independent, and should document where the model must not be used
- Accountability stays with the institution for vendor models
- Change notification clauses are commonly missing and matter
- Rising override rates are an early signal metrics have not caught up to
Frequently Asked Questions (FAQ)
Q: What is model risk management?
It is the discipline of managing the risk that decisions based on models produce adverse outcomes — because the model is wrong, is applied outside its intended purpose, or has degraded since deployment. It covers inventory, validation, monitoring and governance.
Q: What should a model inventory contain?
Every model in use including vendor models, with its purpose and permitted use, named owner, data sources, risk tier, validation date and outcome, monitoring arrangements and last review date. The hardest part is capturing shadow models built informally by teams.
Q: Who should validate a model?
Someone independent of the team that built it. A developer reviewing their own model is not performing validation. Independence is what allows uncomfortable findings to surface.
Q: How do you validate a vendor model you cannot see inside?
Through outcome testing — measuring performance and fairness across segments even without methodology visibility — supported by contractual disclosure rights, challenger model comparison, and a requirement that the vendor notify you when the model changes.
Q: Is the bank responsible if a vendor’s model causes harm?
Under RBI’s FREE-AI framework, accountability stays with the regulated entity regardless of who built the model. This is why institutions are expected to validate third-party models as rigorously as their own rather than relying on vendor assurance.
Q: How should models be prioritised for scrutiny?
By customer impact rather than technical complexity. A simple scorecard declining loan applications warrants more scrutiny than a sophisticated model routing internal documents, because the consequences of error differ enormously.
Q: What is a good early warning that a model has problems?
Rising override rates. When staff increasingly disagree with model recommendations, they have usually observed something before the monitoring metrics detect it. Tracking overrides is cheap and frequently overlooked.
Q: Does RBI require model risk management?
The FREE-AI framework points toward model inventories, validation, ongoing monitoring and board-level governance. As of August 2026 RBI is weighing binding guidelines for banks and NBFCs — verify the current position on rbi.org.in.
Related Reading: