How should we measure quality in SNOMED CT?

SNOMED International has validation tests and KPIs that track SNOMED CT content quality, largely focused on structural aspects such as RF2 compliance, template adherence and avoidance of specifically identified “patterns” such as role group crossovers.

These measures are valuable for ensuring coherence and conformance, but they do not fully address what we might mean by clinical or semantic quality.

In particular, they do not directly test whether:

  • Concept definitions are medically appropriate or clinically valid

  • Parent/child relationships reflect true clinical meaning (not just modelling structure)

  • Content aligns with real-world clinical interpretation and usage

In general those aspects of quality are considered by terminology authors as part of content review, but this approach does not lend itself to quantitative metrics, or anything that could be demonstrated.

This also raises a broader question: are we measuring the right things when we talk about “quality”?

Other domains, particularly software engineering, have developed broader quality frameworks that extend beyond structural validation. For example, standards and approaches such as:

  • ISO/IEC 25010 (SQuaRE model) for software product quality

  • OQuaRE (Ontology Quality Requirements and Evaluation - which adapts software quality dimensions to ontologies)

  • Biomedical ontology evaluation frameworks that separate structural, semantic, and pragmatic quality (eg OEF, OntoQA, SQM)

  • Competency-question based evaluation approaches used in ontology engineering

These approaches consistently highlight a similar gap: structural correctness is necessary, but not sufficient to capture semantic correctness or real-world fitness for purpose.

I’m interested in views on how we might extend or complement current structural KPIs with measures that better reflect clinical appropriateness, semantic validity, and real-world usability. Ideally in a way that is scalable and actionable.

A websearch for work previously done in this area resulted in the list below. I also note that the Member Forum was asked a related question in September here: https://forums.snomed.org/t/preparation-for-member-forum-workshop-content-quality-in-practice/335 (tagging @ahoejen)

  1. Zhang & Bodenreider (2010) — Structural auditing of SNOMED CT using Formal Concept Analysis
    https://pmc.ncbi.nlm.nih.gov/articles/PMC3041382/
    Analyses SNOMED CT hierarchy structure to detect subsumption irregularities and structural anomalies using lattice-based methods.
  2. Mikroyannidi et al. (2012) — Syntactic regularities and irregularities in SNOMED CT
    https://jbiomedsem.biomedcentral.com/articles/10.1186/2041-1480-3-8
    Examines modelling consistency in SNOMED CT by identifying recurring structural patterns and deviations.
  3. Abad-Navarro et al. (2020) — Readability and structural accuracy of SNOMED CT
    https://link.springer.com/article/10.1186/s12911-020-01291-y
    Evaluates lexical readability and structural accuracy of SNOMED CT content across multiple releases.
  4. SNOMED CT usage and mapping evaluation studies (e.g. NLP and coding accuracy assessments)
    https://arxiv.org/abs/2311.10856
    Assesses SNOMED CT quality indirectly via real-world performance in clinical coding and automated mapping tasks.

@jcase I thought this would be a good topic to include at our Joint Advisory Group session planned for October, but I would - of course - welcome any thoughts you have in this area in the meantime.

5 Likes

@pwilliams I couple of months ago we developed a quality assurance document describing the principles and guidelines for ensuring quality of SNOMED CT.

SNOMED CT Quality Assurance Approach, Principles, and Guidelines - v1.5.pdf (273.4 KB)

This includes the use of the existing automated processes, but also emphasizes the knowledge work needed to maintain the terminology.

1 Like

Thanks for raising this, Peter.

I’ve split my response into two parts, to both respond to your question but also highlight opportunities to improve the current processes and hopefully kickstart some discussion…

Measuring quality

Two of the four papers you’ve referenced cover the structural domain that you’ve already noted as insufficient. Detecting non-lattice pairs or axiom pattern irregularities tells us something about internal consistency, but nothing about whether the content is clinically correct.

The third paper I liked. The LSLD metric (“lexically suggest, locally define”) measures the degree to which what a concept is named in natural language is actually represented in its formal logical definition. SNOMED CT scores poorly here, meaning a large proportion of concepts have names that imply attributes or relationships not captured in their stated definitions. This is a directly actionable quality signal, and it has a useful property: it can be applied retrospectively across releases, allowing us to assess whether specific content improvement projects actually moved the needle, and by how much. The caveat is that LSLD is only meaningful where content is sufficiently defined or where the hierarchy has a detailed concept model; which is itself a quality signal about how much of the terminology is actually amenable to formal evaluation at all. A target metric could be determined by first analysing content areas that are considered gold standard, then working to bring the rest of the terminology up to this standard.

The fourth paper is closest to pragmatic quality measurement. Inter-coder agreement of around 75% between trained human coders reflects genuine ambiguity in the terminology, not just training variability. Consistency of real-world application is exactly the kind of complement to structural KPIs that would tell us something useful. I’d have to think more how it might be applied in practice though.

A qualitative signal that is currently underutilised is feedback from the Translation User Group. They surface issues that structural metrics will never detect: terms that are not easily differentiated in English, or that are used interchangeably in clinical practice even where formal definitions differ. This feedback should be informing quality metrics, not just content requests.

The metrics currently used were convenient - they’re easily measured and reported against. But they focus on structure rather than clinical quality - and as a result the Quality work has leaned more in this direction - fixing structural issues rather than true content problems.

The quality process itself

Better measurement is necessary, but it won’t drive improvement unless what’s being measured is connected to meaningful action. This is where I think a more direct conversation is needed.

At CMAG in Oslo, it became apparent that there is a significant gap between what members understand as “Quality Improvement” and what SI has actually been progressing. Members have consistently supported quality initiatives, but the QIP operated with a narrow and specific scope focused on structural conformance (as noted above) - while what most members were hoping for was progress on the backlog of known issues representing real implementation problems. That gap was never adequately recognised until then, and is only starting to improve (albeit, very slowly).

The recently published QA approach document (linked by Jim) is a welcome step, but it remains primarily oriented toward process conformance. Metrics like PPM compliance, template adherence, and RF2 structural rules are tractable and reportable, but they don’t test whether the inferred classification that implementers actually rely on is clinically appropriate. For example: Content can be inconsistently modelled, and still conform to a template.

There are also some structural patterns in how quality work gets prioritised that I think need to be named. There is a sustained preference for smaller, scoped projects over the large complex areas where the real quality debt sits: primitive hierarchies like substances, morphologies, and qualifiers, where improvements would propagate throughout the terminology. Incomplete remediation of a systemic problem can leave content in a worse state than before, replacing coherent incorrectness with inconsistency. And when specific examples raised by members are corrected without addressing the underlying cause, the surrounding problem persists quietly.

The pace of improvement in these areas needs to match the scale of the problem. These issues have not become smaller through deferral, and they will not resolve without deliberate effort at the right level of scope. Better metrics will help make the case – but only if they are tied to improvement targets with real accountability.

SNOMED CT’s value proposition over emerging alternatives rests on deterministic, computable output. That advantage is not self-sustaining: it has to be earned by the quality of the content. As competing tools mature, that value depends increasingly on whether the content is reliably correct and whether we can demonstrate that it is.

I encourage other NRC representatives and implementers to share their experience. A broader signal from the community would be valuable.

3 Likes

Thanks Peter, and I think Matt’s post already covers several of the key points very well, particularly the distinction between structural conformance and clinical or semantic correctness, and the risk that quality work gravitates towards what is easiest to measure.

I would add a few related points.

The first is that the existing issue backlog could itself be treated as an important quality signal. SNOMED International already receives reports of defective content, but those reports vary enormously in scale. Some are isolated concept-level issues. Others point to recurring modelling problems, problematic hierarchies, or areas of known semantic debt. Counting tickets alone would not be enough, but analysing the backlog by domain, scope, age, implementation impact, and resolution status could give a much clearer picture of where the terminology is mature, where it is questionable, and where it is known to be carrying significant debt. Modern LLMs could potentially assist with this kind of backlog analysis and triage.

In particular, it would be useful to distinguish issues that have been resolved from those that have been confirmed but parked, deferred, or made dependent on future projects. Some known problems appear to remain unresolved for years because the larger remediation project has not started or has not delivered. For example there is fairly widespread understanding of the issues in the substance hierarchy and those have been known for many years. That seems like an important quality measure in itself: not just how many issues are raised or closed, but where confirmed defects remain, how old they are, how much content they affect, and what downstream implementation risks they create.

This could support a SNOMED CT content quality heat map. Rather than reducing quality to a single global score, the terminology could be assessed by hierarchy, domain, or modelling area, using signals such as confirmed unresolved issues, age of parked defects, known debt, scope of affected content, implementation impact, modelling maturity, and semantic alignment between the FSN, text definition, OWL axiom, and hierarchy.

Such a heat map could help distinguish areas that are mature and stable from areas that are structurally tidy but semantically under-reviewed, areas with known localised issues, and areas carrying significant pattern-level debt. That would give a much better signal for prioritising quality improvement work than aggregate structural metrics alone. It would also be useful for those using the terminology to better understand the reliability of the areas they might be using and help flag where they need to be careful.

Another related approach is semantic regression testing. We have previously presented the idea of “invariants” for SNOMED CT content development: that an author should be able to record not only the sufficient definition of a concept, but also key things that should remain true about that concept over time. These invariants capture more of what the author knows about the intended meaning of the concept, and provide a regression test if later modelling changes alter the inferred classification or query behaviour.

That presentation is available here: https://www.youtube.com/watch?v=aePzofz5FG4

A closely related idea is to define ECL-based test cases for important clinical use cases: concepts that should always be found, and concepts that should never be found, by a given query. This would make quality more directly testable against the use cases SNOMED CT is expected to support. Together, invariants and use-case-oriented ECL tests could help detect when a content change has unintentionally broken a meaning, classification, or query result that implementers rely on. Extension builders with ECL based reference sets or ValueSets already experience this and it helps to identify issues.

It is also worth noting that alternatives to the proximal primitive modelling style have previously been described which deliver similar error detection benefits. The broader point is that authoring should capture more of the intended semantics than only the sufficient definition, so that unintended changes in classification or query behaviour can be detected earlier.

So, building on Matt’s points, useful quality measures might include:

  • known unresolved semantic debt by hierarchy or domain;
  • age, scope, and implementation impact of confirmed unresolved content defects;
  • whether issues are resolved individually or the underlying pattern is addressed;
  • well known areas with debt requiring refactoring (e.g. substances/products);
  • semantic invariants that should continue to hold over time;
  • ECL-based tests for important clinical use cases, including concepts that should be included and concepts that should be excluded;
  • release-to-release semantic stability, including changes to inferred classification and query results.

The central point for me is that quality metrics should direct effort towards meaningful content improvement. Structural metrics are necessary, but they should not become the proxy for quality itself.

SNOMED CT is valuable because implementers can rely on it for classification, subsumption queries, analytics, decision support, maps, and interoperability. That requires more than structural regularity. It requires confidence that the stated and inferred meanings are actually correct, and that known areas of semantic debt are visible, prioritised, and progressively addressed.

This is a highly appreciated initiative! I agree with the previous comments, but would like to highlight SNOMED CT as an interface terminology. When translating SNOMED CT into other languages, the measuring the quality of the translation process and output is important. Quite often this is contingent on the quality of the source language, i.e. the international edition. A recurrent issue is linguistic consistency, also called natural language consistency. An example is whether the description of History of X matches the description of Disorder X. There is an interesting paper which could be added to @pwilliams list above: “A model for Evaluating Interface Terminologies” by Rosenbloom et. al (2008). This papers adresses several parameters which is quite relevant when discussing quality.

2 Likes

Really good point @ovage, @mcordell and I had discussed this too - translation is a great way to find content quality problems.