AI Model Risk vs System Risk: What a Model Score Misses
Severity is not risk: a model risk assessment can't give you your EU AI Act tier or your deployment risk, and here is what to demand instead.

"Benchmarks don't catch this. Risk assessment does." That was the line on a recent LatticeFlow AI post, set over a screenshot of a model graded like a vulnerability report: Open-ended Correctness 7.7 (High), Hallucinations 8.4 (High), Prompt Injection 5.2 (Medium), each score assigned "as per CVSS v4.0 severity levels." The line is half right, because benchmarks genuinely do not catch deployment risk. The difficulty is that a model risk assessment does not catch it either, and for the same reason: both of them are, at heart, measuring the model, and a model is not what you deploy; even a score annotated with a few system descriptors is not the use-based assessment of the whole system you actually run, which is where the risk a governance team has to answer for lives. LatticeFlow is one of a growing list of vendors now selling a model-level risk score, and the argument here is about the shared approach across the category. What follows sets out what a model risk assessment genuinely tells you, where that evidence stops, and what a governance buyer should ask for instead.
A whole category now scores the model
Model-level risk scoring has become a genre, and at its best it is real and useful work. LatticeFlow AI's AI Insights service, and the open-source COMPL-AI framework it co-developed with ETH Zurich and INSAIT, benchmark models on security, fairness, and EU AI Act-aligned criteria; Lakera's AI Model Risk Index scores models for adversarial resilience through its b3 benchmark; Enkrypt AI publishes an LLM Safety Leaderboard with a single risk score spanning jailbreak susceptibility, bias, toxicity and malware generation; and Cisco AI Defense, built in part on Robust Intelligence's AI security work, offers model and application validation and red-teaming. Patronus AI's evaluator models and the academic Aymara LLM Risk and Responsibility Matrix sit alongside them in the broader model-evaluation family.
These evaluations earn their place: they help you choose between models, catch gross failures before anything ships, and generate monitoring signals, and in one specific case they can contribute to satisfying a legal duty, since the EU AI Act does require model-level evaluations of general-purpose AI models that carry systemic risk. The better vendors are candid about the boundary, too: Lakera says plainly that its benchmark tells you which models are most resilient while a separate red-team tells you whether your system is secure, and LatticeFlow's assessment ingests system-context factors while the COMPL-AI framework states openly that it is "not official auditing software for EU AI Act compliance." The problem begins when a model-level severity score is treated as a system-level governance verdict, which is the category error at issue.
What CVSS severity actually measures
Look again at the small print under that table: the scores are assigned "as per CVSS v4.0 severity levels." CVSS, the Common Vulnerability Scoring System, is the standard the security world uses to rate software vulnerabilities, and it is a precise, carefully built instrument for what it does. What it does is measure severity, and its own maintainers at FIRST are emphatic that severity is a different thing from risk, since their guidance states that a CVSS Base score is "designed to measure the severity of a vulnerability and should not be used alone to assess risk," because that base score captures only the intrinsic characteristics of a vulnerability, independent of the environment it sits in and the threats acting on it. You only get close to risk once you put context around that intrinsic number, the environment the vulnerability actually sits in and the threats actually acting on it.
Borrowing the CVSS severity ladder to grade a model imports the warning along with the scale. A "Hallucinations 8.4" describes a model property on a vulnerability-severity scale, and by itself it tells a business neither what it stands to lose nor which obligations it triggers; it starts to speak to risk only once two kinds of context are added, the intended use the system is put to and the deployed system around the model.
Your EU AI Act risk is set by use
Start with the regulation, because it settles the question cleanly. The EU AI Act classifies AI by intended purpose, defined in Article 3(12) as the use for which the provider intends a system, including the specific context and conditions of that use, and the obligations follow from it. Prohibited practices in Article 5 are written as uses, the placing on the market or use of a system for a banned purpose such as social scoring; the high-risk category in Article 6 is reached either through the product-safety route of Article 6(1) or by falling into one of the use-cases listed in Annex III under Article 6(2), employment, creditworthiness, biometrics, access to essential services and the rest, subject to the narrow derogation in Article 6(3), which is never available for profiling; and the transparency duties in Article 50 attach to how a system interacts with people or generates content. Not one of these triggers is a model property.
In practice, a system that screens job applicants with a near-flawless, low-hallucination model is high-risk under Annex III, with the full weight of conformity assessment, logging, human oversight and post-market monitoring attached to it, while a system that uses a hallucination-prone model for internal brainstorming falls outside the high-risk regime, with at most the ordinary transparency duties that apply when a system interacts with people directly. The model score moved neither system into its tier; the intended use did.
The one place where the Act does reach the model itself is Chapter V: every provider of a general-purpose AI model carries documentation and transparency duties there, and providers of models with systemic risk carry evaluation and mitigation obligations on top, triggered by high-impact capability, with cumulative training compute above 10^25 floating-point operations creating the presumption; a hallucination rate or a prompt-injection score has no part in that trigger, and the regime stays separate from the high-risk classification of the systems you deploy.
The same model, two different risks
The second context a model score cannot resolve is the deployed system. A model evaluation measures the model under standardized test scenarios, and that is not how anyone runs it in production, where what you actually deploy is a system, the model wrapped in a harness, a system prompt, retrieval, tools and guardrails, with risks that have little to do with the model's intrinsic scores. Take one model and deploy it twice: in the first system it sits behind retrieval grounding, a constrained output format, and a person who reviews every result before it reaches anyone, while in the second it is handed tool access and the autonomy to act, reading from and writing to live systems with no one in the loop. The model is identical, its benchmark scores are identical, and the real risk is nowhere near the same, because risk is a property of the whole sociotechnical system, and the weights are only one part of it.

NIST frames it the same way in the AI Risk Management Framework, where risk emerges from the interplay of the technology with how a system is used, who operates it, and the context it runs in; its Map function exists to establish that context before measurement. The evaluation-research community reaches the same conclusion from the technical side, where work on sociotechnical safety evaluation argues that context decides whether a given capability can cause harm at all, which leaves a capability measured in isolation unable to establish the harm.
The clearest confirmation comes from inside the category: the COMPL-AI framework itself concedes that human oversight and corrigibility are system-level concerns it cannot evaluate at the model level, and that rigorous tools for measuring explainability do not yet exist, and those concessions sit precisely where the EU AI Act's high-risk obligations concentrate.

Answering the dashboard critique
LatticeFlow makes a sharper argument against governance platforms, and it runs roughly like this: a compliance dashboard rates how good your paperwork and workflows are while saying nothing about whether your AI systems behave safely in the real world. As a criticism of governance theater, it lands, because a registry full of green checkmarks that no one has tested against a model's behavior is its own kind of fiction.
The conclusion does not follow, because a model score with no use-and-system frame is the mirror image of the same failure, a precise number measuring the wrong unit. You need both: the technical evidence about how a model behaves, and the use-based classification and system-level assessment that turn that evidence into a position you can defend in front of a regulator. A model evaluation feeds governance, and it cannot replace it; selling it as the whole answer repeats the category error from the other direction.
What to actually demand
If you are buying or building this, the model score belongs in the file as evidence, while four things carry the actual weight. Classify every AI system by its intended use and run it against Annex III, because the use-based tier is what sets your obligations, and no benchmark score moves a system in or out of it. Assess the deployed system, its context, human oversight, degree of autonomy, data and downstream use, and quantify that risk in monetary terms across the whole system, so what reaches the board is an expected loss it can set against budgets and risk appetite, which a letter grade never gives it. Keep model evaluations for what they do well, model selection, monitoring, and the model-level duty around systemic-risk general-purpose AI, and map the evidence once across the EU AI Act, NIST AI RMF and ISO 42001, so a single piece of work counts in more than one place.
Modulos works at that system layer: it classifies each AI system by intended use, assesses it in its deployment context, and quantifies the resulting risk in monetary terms for the system as a whole, a figure a board can act on instead of a "High" on an undefined scale, while holding the model evaluations you run as evidence in a Governance Graph mapped to the EU AI Act, NIST AI RMF and ISO 42001, so a model score sits where it belongs, as one input to a defensible system-level position. To see that structure across your AI inventory, request a demo and we will walk through the classification, the assessment and the evidence map.
Benchmarks do not catch deployment risk, and a model risk score does not either, because what you ship is never the model alone. The risk you answer for belongs to the system that puts the model to use, and the unit that turns it into a board decision is money.
Frequently asked questions
Does a model risk assessment tell me my EU AI Act risk classification? No. The EU AI Act classifies an AI system by its intended purpose and use-case (Article 5 prohibited practices, Article 6 with Annex III for high-risk, Article 50 for transparency), so the tier follows from what the system is used for, and a model's hallucination or prompt-injection score does not affect it. The model-level regime in Chapter V runs separately: it covers general-purpose AI models as such, with additional duties for those with systemic risk on a capability trigger and a compute presumption.
Is a CVSS-style score an AI risk score? No. CVSS measures the severity of a software vulnerability, and FIRST, which maintains it, states that a base score "should not be used alone to assess risk." A CVSS-style model score reports severity, and it speaks to risk only once the intended use and the deployed system are taken into account.
Does the same model have the same risk in every deployment? No. The same model can be low-risk behind retrieval grounding and human review and high-risk wired into an autonomous agent with tool access. NIST's AI Risk Management Framework treats risk as a property of the sociotechnical system in its context of use, with the model only one part of it.
Are model evaluations useless, then? No. They are valuable for model selection, for monitoring, and for the EU AI Act's model-level duty on systemic-risk general-purpose AI; what they cannot do is classify your deployed system or quantify its exposure, which needs the surrounding assessment.
How should AI risk be measured instead? At the system level, in terms a business can act on: the EU AI Act tier that sets your obligations, and the risk quantified in monetary terms across the deployed system, so you get an expected exposure in money instead of a "High" on an undefined scale.
Ready to Transform Your AI Governance?
Discover how Modulos can help your organization build compliant and trustworthy AI systems.


