On the use of observables in concept definitions

Dear MAG

An addition to the September 2026 Editorial Guide (‘Findings and Observables’) caught my eye. It states:

“The term observations should not be confused with Observable entity. Observable entity is the name of something that can be observed and represents a question or assessment (e.g. |systolic blood pressure|, |color of iris|, |gender|) which can produce an answer or result.

Given differences in information models, a finding about the subject of the record may be captured in different ways. For example, the information may be captured using either an Observable entity concept together with a value, or a Clinical finding concept which represents what is being observed together with the result of the observation. To assist the transformation between these concepts, if when adding a requested Clinical finding concept, the modeling requires an Interprets relationship, and the Observable entity concept does not exist, then a new Observable entity concept should be created.”

I understand the practical motivation behind this guidance, but I think it exposes a longstanding modelling problem (and fundamental concept representation problem) that deserves further explanation before creating additional content in line with this guidance clause becomes routine. This is a topic on which I have written before (e.g. here and here) but am yet to receive a convincing or reassuring answer.

For observations associated with ratio or interval quantities, the ‘Observable + value’ pattern is generally justified and understandable. An observable such as mass concentration of sodium in serum includes a recognisable property, and there is a reasonably clear and reproducible boundary between the observable and the value/result such as 140 mmol/L.

The position is much less clear where the finding takes a nominal value, a situation made more complex (a) where the created observable has no readily understandable property type and (b) where the same ideas are handled differently in SNOMED’s formalism.

What does the new observable represent?

The part of the new guidance that concerns me most is:

“…then a new Observable entity concept should be created.”

What rules determine the nature of that observable? In particular:

  • Should it include reference to a recognised property type (and if not, what then is being ‘observed’)?
  • Should its expected value datatype or permitted value set be declared (see below)?
  • What governs the allocation of semantics between the observable and its value?
  • What evidence demonstrates that independent authors would create the same observable for the same purpose?
  • How should its equivalence with the corresponding Clinical finding be established (and/or what does the vague “…assist the transformation between the concepts…” mean in practice)?

Without answers to these questions, the new guidance appears to permit almost anything that can be framed as a question to be reverse-engineered into an Observable entity.

The existing Interprets/Has interpretation mechanism provides only partial support here. A Clinical finding may be defined in terms of an ‘Observable + value’, but nothing guides whether the selected partition is appropriate or reproducible. The value set for Has interpretation has grown considerably in recent years (a near 20x increase from 2021 (200) to 2026 (3800)) but still there are no visible rules for determining the values suitable for specific Observables or corresponding property types. The problem is compounded further where semantically proximate findings use different modelling patterns, such as Finding site/Associated morphology (compare, for example, the modelling approaches to 246679005 |Discharge from eye| and 248952001 |Discharge from female genitalia|, the latter defined using an observable added in September 2026).

Models of capture are not necessarily models of meaning

There is also a broader issue arising from the increasing use of an ‘Observable + value’ content defining pattern.

A structured assessment is a legitimate model of capture – I understand and support such an approach if done well (see policy statement 4 here for an ACP position on structured recording). Where a recognised instrument asks a defined question, constrains its permitted answers and records the relevant context, an ‘Observable + value’ recording structure accurately preserves both how the information was elicited and the resultant semantics.

However, it does not follow that the ‘question + answer’ structure should automatically become the canonical model of meaning (as will be the case for many SNOMED CT Findings defined using the Interprets/Has interpretation modelling pattern).

Clinical records will also contain direct, nominalised assertions (as Findings). A clinician may record that a patient has green eyes, is at high risk of suicide or has a discharge from a particular part of their body. No explicit ‘question’ may have been asked and no separate ‘answer’ given.

These two (incompletely comparable) patterns will become increasingly relevant to the coding steps used with ambient voice technology and other AI-assisted documentation. A system receiving a narrative statement may suggest codes based on:

  • mapping it to a Clinical finding representing the assertion; or
  • decomposing it into an Observable entity and nominal value.

Where the (latter) decomposition is not governed by reproducible rules, different systems may make different (and analytically consequential) choices. Editorial variability will then begin to appear in clinical coding workflows, with the clinician author or coding algorithm arbitrarily deciding where the observable/value boundary should lie.

The code suggestion workflow may also add a potential review burden/friction. A clinician may be expected to review a suggested code for ‘green eyes’ to assess whether it reflects what was said during the encounter, before committing it to the record. Reviewing multiple codes presented as ‘eye colour’ (or ‘iris colour’) plus a ‘green’ value requires the clinician to validate a decomposition presented as a ‘question that was never asked’ and an ‘answer that was never given’. This may be structurally elegant, but it is not necessarily the most transparent representation of the source assertion.

Additional points

A similar sentiment to the ‘Findings and Observables’ section above is made in the Cancer Synoptic Reporting Implementation Guide. This section helpfully describes the sorts of content that would be needed to serve as ‘answers’ to the synoptic report questions (“…values from the Procedure hierarchy, Body structure hierarchy, and concepts in the qualifier hierarchy NOT subsumed by << 260245000 |Finding value (qualifier value)|…”). The use of such values in record instances will likely introduce even more opportunity for the same clinical ideas to be represented in non-isomorphic (and thus analytically remote) ways (consider, for example, one record containing 1660001000004100 |Histologic type of primary malignant neoplasm of breast| and a value of 863926008 |Angiosarcoma| and another record containing 721576006 |Primary angiosarcoma of breast|. I’m not saying they couldn’t both be detected, but…).

Conclusion

To me the implication of both these guidance document sections is that SNOMED CT now explicitly supports - by design - incompletely constrained and logically incomparable mechanisms for recording data items that are intended to mean the same thing (at odds with Cimino’s non-redundancy desideratum). The documentation is welcome, the implications are concerning.

Such an approach may be a necessary compromise in order to meet multiple use cases. However, if the SNOMED community wishes to continue down this route then my suggestions would be:

  • Content created using an ‘Observable + value’ definition pattern needs to be managed with more care to limit variation - in particular in the design and selection of suitable ‘observables’, and with consistent application of modelling patterns to semantically proximate content.
  • The user community - operating at all points in the data lifecycle - needs to be fully aware of this design philosophy and the risks that exist where such an approach is used
  • The consequences and risks of this approach need to be explicitly managed to minimise the likelihood of, for example, false negatives in data retrieval and arbitrary patterns of code suggestion in NLP and AI-based coding workflows

I would welcome the MAG’s thoughts.

Kind regards

Ed

1 Like