2023-04-03 - TRAG Meeting Agenda/Minutes

2023-04-03 - TRAG Meeting Agenda/Minutes

Date

  • Monday 3rd April 2023  -  13:30 - 17:00 (GMT) (12:30 - 16:00 UTC)

Room: Fenchurch

 

  • Wednesday 5th April 2023  -  13:30 - 16:30 (GMT) (12:30 - 15:30 UTC)

Room: Fenchurch

 

Dial in details:

 

Attendees

  • @Andrew Atkinson, Chair

  • @Mounir Bouzanih (Unlicensed), member

  • @Mikael Nyström, member

  • @Patrick McLaughlin, member

  • @Alejandro Lopez Osornio, member

  • @Stuart Abbott (Unlicensed), member

  • @Matt Cordell, member

  • @Gábor Nagy , member

  • @Dion McMurtrie, guest/observer

  • @michael lawley, guest/observer

  • @Reuben Daniels, guest/observer

  • @Chris Morris, staff/observerst

  • @Maria Braithwaite , staff/observerst

  • @Janice Spence Observer

Apologies

  • @Harold Solbrig, member


Objectives

  • Briefly discuss each item

  • Agree on the plan to analyse and resolve each issue, and document the Action points

  • All those with Action points assigned to them to agree to complete them before the next face to face conference meeting

Discussion items

 

 

Subject

Owner

Notes & Actions

 

Subject

Owner

Notes & Actions

1

Welcome!

All

Thanks to our members for all of their help. Welcome to our observers!

INTRODUCTIONS...

We've got several topics that we've resolved and closed down

As always, we won't waste time going through them again in detail, but if you'd like to read through them they're listed below...  

I'll also run through them very quickly from a high level, and if you have any further questions/news on any of the discussions please let me know now and we can decide whether or not to re-open them...

2

Conclusion of previous Discussions topics

3

MedDRA Production release

 

The first SNOMED CT MedDRA Simple Map package Production Release will published on 30/04/2021

This will include 2 maps - full details will be included in the Release Notes.

Does anyone have any last minute questions/issues to raise before the Production release is published?
No - in which case we'll proceed as planned
Topic to be closed down in October 2021 TRAG meetings (after Production release) unless new issues are raised...
 
TOPIC TO BE RE-OPENED DUE TO FEEDBACK FROM THE COMMUNITY ON THE FORMAT OF THE MedDRA to SNOMED MAP FILES:
We should open up a new topic to review the proposal (incoming from the Implementation team) for a new format for reverse direction maps...
 
FINAL DECISIONS:
 
This has been agreed in the topic "Redesign of the Map Reference Set formats"
We will take this proposal to the MAG, and if ratified we will:
a) take the plans forward with the content team, in order to include the necessary new concepts in the January 2022 International Edition
b) update the RF2 spec accordingly
 
NOW we need to agree how to communicate it out to the community ahead of the impending 2022 MedDRA release...
Is it enough to:
a) send out general comms to the Release distribution list confirming the upcoming changes
b) + send the same comms out to those users who we know downloaded the April 2021 MedDRA release package?
Or do we need to do something more?
No, that's adequate 
 
In addition, we need to agree how to build the MedDRA package in 2022, in order to clearly show a distinction from the April 2021 release (in the old format), whilst also retaining the historical audit trail.
Everyone agreed that we need to produce the April 2022 MedDRA package:
a)  in the new format (as per "Redesign of the Map Reference Set formats")
b)  with the April 2021 map file(s) removed
c)  BUT the new format map files should contain both the new data, PLUS all the historical MedDRA data (from 2021) in the NEW FORMAT.  This means that the NEW file should look exactly like it would be if we had actually published the original April 2021 MedDRA release in the NEW FORMAT (with all original data from 2021 + all new inactivations/changes from the latest cycle)
 
NEW RELEASE PACKAGE IN THE NEW FORMAT HAS NOW BEEN PUBLISHED IN THE 2022 PRODUCTION MEDDRA RELEASE
ANY FURTHER FEEDBACK??
HAVE WE NOW RESOLVED ALL KNOWN ISSUES AND CAN CLOSE THIS TOPIC DOWN???

4

The possibility of updating inactive content

All

https://projects.jira.snomed.org/browse/MSSP-1670

Please see the ticket above for full explanation - in brief:

  • Descriptions in the 20220731 International edition snapshot description file appear to contain ASCII Character 160 for Non-breaking space, when the character should be ASCII 32 for a standard space. ASCII Character 160 could potentially create issues with ETL processes.

  • The suggestion is that issues caused by non-printable ASCII / UNICODE / UTF-8 characters need to be covered under their own policy because simple inactivation does not resolve the issues caused by these characters in ETL and interoperability processes.

  • Unfortunately removing these characters from inactive content contravenes our current policy, which is to only update inactive content (whether this be via the AP or via a back-end fix by the tech team) where a "critical issue" has been found. The term "critical" is used specifically to clearly denote only those issues which present risks such as clinical patient-safety or legal liability, for example.  Therefore, in order for us to flag up these inactive records as a clinical safety issue, we'd need evidence of reports from users explaining how they present such a risk to their patients.  

  • Confirmed by the content team that validation for non-breaking spaces is in place already for active content, and so no improvement to validation is required.

  • From what we can tell, many of these have been in the release for years now, but we have not received any feedback that it has caused an issue thus far. This is therefore not a "critical" issue - however we'd appreciate community to confirm if there would be any issues with making the fixes directly on inactive descriptions?

  • The following Descriptions in the 20220731 International edition snapshot description file were found to contain ASCII Character 160 for Non-breaking space, when the character should be ASCII 32 for a standard space. ASCII Character 160 creates issues with many ETL processes.

  • All of these issue are in inactive descriptions:

  •  

    ID Column Issue
    2869833013 [term] Code:160, Position:27
    2870804019 [term] Code:160, Position:27
    2871691013 [term] Code:160, Position:16
    2880511019 [term] Code:160, Position:21
    2880958016 [term] Code:160, Position:118|Code:160, Position:149
    2881152012 [term] Code:160, Position:118|Code:160, Position:124
    2882107012 [term] Code:160, Position:118|Code:160, Position:149
    2882999016 [term] Code:160, Position:118|Code:160, Position:124
    2884068015 [term] Code:160, Position:21
    3030804017 [term] Code:160, Position:30
    3030901012 [term] Code:160, Position:30

  • The suggestion is that "issues caused by non-printable ASCII / UNICODE / UTF-8 characters need to be covered under their own policy because simple inactivation does not resolve the issues caused by these characters in ETL and interoperability processes. Given the amount of inactive SNOMED content present in the data stream, it would be best if these characters could be removed entirely from even inactive descriptions. While working in healthcare implementations, the presence of ASCII Character 160 (non-breaking space) in the LOINC descriptions broke the entire ETL process between the data warehouse and Research databases and required me to jump through some programming hoops to remove these characters from the LOINC descriptions."

  • Whilst we appreciate the impact that these characters might have on ETL processes,  from a content team perspective, this is not a critical issue.  All of the current issues are related to inactive descriptions and validations are in place to prevent this from occurring in the future. 

  • Unfortunately removing these characters from inactive content contravenes current SI policy, which is to only update inactive content (whether this be via the AP or via a back-end fix by the tech team) where a "critical issue" has been found.  The term "critical" is used specifically to clearly denote only those issues which present risks such as clinical patient-safety or legal liability, for example.  Therefore, in order for us to flag up these inactive records as a clinical safety issue, we'd need evidence of reports from users explaining how they present such a risk to their patients.  

  • SI are always reluctant to change SNOMED CT history. However, there are situations where we have had to do that in the past. We are therefore bringing this to the TRAG for consideration...

  • A couple of people thought it might be easier to update the inactive content rather than getting repeated complaints over the years - however the vast majority disagreed, and thought that not only was it a waste of valuable resource to update inactive content, but more importantly actually contravened the spec at this level!  This is because the INT Edition specifies itself as a UTF-8 format, and the ASCII 160 characters are UTF-8 compliant!  Therefore where would we stop once we start excluding certain UTF-8 characters from the INT Edition? 

  • Instead, it should be the responsibility of the end implementations to exclude any characters that disagree with their ETL routines/programs.

  • Repsonse added to https://projects.jira.snomed.org/browse/MSSP-1670

  • TRAG RECOMMENDATION WAS TAKEN - our SI specs specify UTF-8 format, and as ASCII 160 characters in question are UTF-8 compliant, any changes would be contravening our own specifications.  Therefore all were agreed that no changes should be made in the content - instead it should be the responsibility of the end implementations to exclude any UTF-8 characters that cause issues for their ETL routines.

5

AttributeValue field immutability in the RF2 files

ALL

Just a very quick one (especially for those who were in the MAG yesterday and have already heard this!) - the immutability of the valueID field is specified as being "depends on specific use" - see here:

The MAG are all happy to change this to "mutable", and so are we - however I just wanted to give those here who weren't in the MAG a chance to raise a valid objection in case anyone can identify a really strong reason why this field shouldn't be mutable??

No objections raised

6

Active Discussions for April 2023

 

7

Welcome and thank you!

 

Welcome to new members!

8

Member Nominations

 

Please let us know if anyone is interested (and who has the requisite domain knowledge and expertise) in applying for a seat on the TRAG - thanks!

We're looking for new members to take the place of some outgoing chairs - if you have any Nominations please let me know either this week or by email after.   Thanks!

9

Derivative product Release package formats

Kai

The SI Standards for Derivative packaging formats (and therefore the precedents set for all existing Releases such as MedDRA, GMDN, etc), are that we create packages that are inherently dependent on the relevant International Edition package (as per the MDRS). 

These Derivatives are solely single refsets/maps, and so don’t mean anything to the end users without the supporting terms and other components from the International content. 

This is why we have always created the metadata components (refset/module concepts, descriptions, relationships, etc) in the International content, and then made the Derivatives dependent on the relevant International Edition.

Recently however, during the creation of the new EDQM maps, the question was raised as to whether or not we should change to include the metadata concepts in Derivative packages themselves, rather than in the International Edition.  This would not make them stand-alone (like the IPS sub-ontology for example), as they would still be dependent on the International Edition.

However, as a Derivative with its own module there may be benefits of the package containing its own metadata concepts?

The main benefit so far identified has been to avoid the situation experienced in the EDQM Alpha release:

  • the metadata for the EDQM product was published in the December 2022 International Edition,

  • and so the EDQM map content was upgraded in line with the December 2022 International content.   

  • the maps themselves however were not updated, and were still based on July 2022,

  • but the release itself was dependent on the December 2022 International Edition.  

  • this created a slight contradiction in terms of the map release which appeared to be based on December 2022, but actually the metadata itself was the only thing baselined on this release.  The content was only brought up to date with the July 2022 release.

  • Having said all this, the flip side is that if we just push the collaborators to request the metadata to be added to the International Edition at an earlier time, then this situation can also be avoided!

HOWEVER, is it possible that the issue below in the RT2 release process could also be mitigated by moving the metadata into the Derivative packages themselves?  If we did this, could we potentially then retain the Derivative content in the Refsets' child branches and release from there?

(eg) 

  • For GPFP we would maintain the content via the RT2 FE which would read/write to MAIN/Refsets/GPFP

  • Given that we would then also include the Module + other metadata in the same branch (instead of in MAIN),

  • ...Could we then version say into MAIN/Refsets/GPFP/2023-04-30, and Release straight out of that branch??

 

Changing the way we do things currently would take some work, and so would require a strong business case in order to 

a) Change our standards

b) Incur the cost of making the technical + release changes

Thoughts on the original decision to package them in this way?

Everyone comfortable that this was the right decision at the time, but not anymore.

Thoughts on the benefits/risks of refining this standard?

Consensus is that moving the Metadata components into the Derivative packages brings benefits to both the maintainers (SI + NRC's) + also to the end users, as it's easier than having to pull the metadata down from the dependent International Edition.
It doesn't have a huge impact however, as the end users still need to download and consume the relevant International Edition when using the derivatives.
The big benefit however, comes from combining this change with the new way of managing refsets in RT2 (see below for full details).  The inclusion of this metadata in the Derivative branches in Snowstorm (instead of in the INT branch)allows us to move the Derivatives into their own codesystems, and thereby allow us to retain the different dependencies of the Derivative products on older version of the INT Edition than it would do otherwise in the new world of RT2
(as things currently stand the Derivatives would have to be rebased against the LATEST INT Edition content whenever we need to promote them up to MAIN to publish them.  This would force us to bring the Derivatives up to date with the very latest INT Edition content just before publishing the Production Derivative releases, which would be April/October.  All users have confirmed this would be a huge problem for them as they're not yet ready to take multiple monthly releases of the INT Edition, and so desperately want us to retain the dependencies on the Jan/July INT Releases) 
Therefore we can continue to Publish the Derivatives in April/October (which is the earliest possible date for most of them because of the delays in collaborating with the external entities), based on the Jan/July INT Edition releases, exactly as the users want.
 
If we agree on moving the module concepts to the Derivative package(s), do we also need to remove them from the International Edition, in order to avoid duplicates when users implement them in conjunction with the International content?
  •  

    YES - we need to put a proposal forward to:
    a) DEMOTE the various Derivative metadata components down from the INT Edition in the July 2023 Release
         (this would simply involve inactivating them all in the International Edition)
    b) PROMOTE  them in the various Derivative packages in the Sept/October 2023 Derivative releases.
         (this would simply involve activating them in the relevant Derivative packages)
    ********* SEE OPTION 2 in the Refset  processes in the post-RT2 agenda item for FULL proposal ******
     

10

Refset  processes in the post-RT2 world

All

Options to be discussed (see local Notes "RT2 New Process")

MAIN POINTS:

  • The International Derivative releases have always been dependent on Jan/July, in order to align with the fact that most end users still only take the Jan/July releases, and not the other monthly releases.

  • The new RT2 integration with snowstorm branches greatly improves the quality, because it allows the authors to run full validation against the new derivative content BEFORE promoting, thereby allowing identification + fixing of most issues before they hit the Release cycle. (eg)

    • MAIN/Refsets/GPFP,   MAIN/Refsets/ICNP,  etc

  • However, in order to actually RELEASE these derivative products, we’ll need to Promote the Refset content up to MAIN/ (as the metadata content is in the International Edition, and we can’t currently export content from multiple snowstorm branches into SRS for the same release, just from one branch)

  • Before you’re allowed to promote in the AP, you HAVE TO re-base and validation FIRST

    • The re-base process will pull the content through from the LATEST International content in MAIN, which may include monthly content from later releases than the recent JAN/JULY release that we originally wanted to base the Derivative release upon!

  • This makes it inherently less flexible in terms of preventing re-basing against later monthly International releases, and therefore extremely difficult to retain dependency on JAN/JULY releases.

OPTIONS:

  1. BEST option appears to be to allow the Derivative refsets to be dependent on the month prior to that in which they're Published (eg) March/September for most derivatives.

  2. However, is it possible that if we were to make a different decision on the item above, then this issue in the RT2 release process could also be mitigated by moving the metadata into the Derivative packages themselves?  If we did this, could we potentially then retain the Derivative content in the Refsets' child branches and release from there?

    (eg) 

    • For GPFP we would maintain the content via the RT2 FE which would read/write to MAIN/Refsets/GPFP

    • Given that we would then also include the Module + other metadata in the same branch (instead of in MAIN),

    • ...Could we then version say into MAIN/Refsets/GPFP/2023-04-30, and Release straight out of that branch??

  3. The final option is moving to Monthly releases for Derivative products as well - however:

    1. This is problematic from a capacity/practical point of view, and

    2. This would require significant discussions with all of our collaborating entities to sign off on new agreements (for those who have tangible inputs into the process), including monthly deadlines, etc - This could not be achieved quickly so would need to be discussed with 2024 in mind at the earliest. 

      1. The alternative to this would be to move just SOME of the derivatives (the straightforward refsets for example which don't require much/any external collaboration) to monthly schedules, and keep the others on annual/6 monthly schedules.

      2. However, this would be a significant operational issue, as the processes are already complex enough without introducing further complexity in terms of different derivatives runing on completely different processes.

We should discuss options and agree the best way forward for retaining quality within the Release process vs impact to the users.

Option 1
No-one is in agreement with this option!
This would cause real problems for NRC's (let alone end users) who would struggle to keep up with downloading and consuming multiple monthly releases.  In addition, they wouldn't be able to then publish multiple releases of their own to the end users, containing the usual Jan/July changes + then extra releases for each month that is a dependency for the Derivative releases.
Finally, many of the entities doing Translations are struggling enough to keep up with the pace, without adding additional stress and complexity.  This option would therefore completely prevent them from consuming any of the Derivative products.
This Option is a non-starter.
 
Option 2 
The Pro's are far greater here - as all of the users who can't keep up can continue to download just the Jan/July INT Edition releases, and then consume whichever Derivatives they need.
Therefore we need to put forward this proposal, which would be to:
a) MOVE all of the Derivative content in Snowstorm from the INT Edition branch into their own Codesystems (checked with Rory and Terance and should not be a problem)
b) CHANGE RT2 to use the new individual codesystem branches for reading + writing each Derivative content to and from RT2 (checked with Brian and Rick and should not be a problem)
c) DEMOTE the various Derivative metadata components down from the INT Edition in the July 2023 Release
     (this would simply involve inactivating them all in the International Edition)
d) PROMOTE  them in the various Derivative packages in the Sept/October 2023 Derivative releases.
     (this would simply involve activating them in the relevant Derivative packages)
e) THEN EACH CYCLE WE WOULD:
i)  Upgrade each Derivative Codesystem to the relevant Jan/July INT content
ii) FREEZE the content in each of those "in flight" codesystems, to prevent any more re-basing until the Release cycle is complete
iii) VERSION in each relevant Derivative Codesystem
iv) RELEASE from each relevant Derivative Codesystem
 

11

Community Consulation: Proposed changes to the RF2 Identifier File Specification

All

Full details can be found in:  SNOMED International Proposal to change the RF2 Identifier File specification

Main points for TRAG consideration:

  • Proposed Changes to RF2 - any issues with the proposal?

    The current column headers for this file are:

    The new proposed format is:

    The data types will remain the same, as detailed in the current RF2 specification:  4.2.4 Identifier File Specification   

  • Perceived Impact - is this the case?

    This change is not expected to have any impact on implementers of existing systems, as SNOMED International are not aware of organisations who currently represent entities from other code systems directly in SNOMED CT, as opposed to mapping to it.   As such, consumption of the new file would only be required by organisations who have an interest is working with such content.

  • FEEDBACK - Walk through the feedback received so far and confirm if any further points?

Feedback requested:

Feedback on the File changes was varied, but generally speaking there were no strong objections to the changes to the file.
HOWEVER, there were strong objections to the overrall plan to publish LOINC as a separate "Extension".
This is due to the additional Friction caused by having yet another component in a separate package.
Implementers would greatly prefer it to all be published in the same package as the International content.
Having spoken to Rory, this is a CONTRACTUAL issue - we cannot align the SNOMED CT licence with the LOINC licence in order to publish both types of content in the same package!
This is therefore the ONLY option - we have to publish both the LOINC Identifier file + the LOINC content itself in a separate Extension package, dependent on the International Edition.
 
So we now need to go back to the intended changes to the Identifier file format, and confirm whether or not these are acceptable to everyone?
Initial feedback:
We're making it "look" like the other RF2 files, but it's not!  The Identifier column is NOT a primary key as you'd expect, as in other files with UUID's (even though they also technically have compound keys such as UUID+moduleID+active, etc)
We'd have to make the "ReferencedComponentID" field mutable, as otherwise when a mistake is made and we need to change this field to another ID, we have no other option than to create a DUPLICATE record which has everything the same except for active+ReferencedComponentID.
This shouldn't be too much of a problem though as we can make the ReferenceComponentID field mutable if we need to
Most people would prefer to use a Refset instead in order to be more flexible
We could have a unique primary key (like a UUID)
We could express one-one and one-many relationships, etc
URI attributes such as Concrete Domains coudl be a much more useful addition to the identifier file?
.
 

12

IPS Terminology Product

All

Quick run through of the changes that we're proposing to make in the final Production release in Q4 2022, as compared to the BETA release

(ie) discussion of the feedback that we accepted and have implemented in the Production release:

  1.  INCLUSION OF THE “EML” (new Drugs refset) IN THE FEEDER FOR THIS PRODUCT FROM 2022 ONWARDS

  2. IPS Terminology URI:

         *** PLEASE SEE SECTION E here for final solution:

ACTIONS:

Reminder that this is a SNOMED International product, but NOT a SNOMED CT product, which means it's non conformant to many of our normal standards

Another reminder that this product is NOT for members, it's only useful for non-members (mostly those new to SNOMED)
Questions on any changes planned?

OCTOBER 2022: Any final feedback before finalise first Production release?

AAT to discuss with the business and come back to everyone with potential solutions on Wednesday...

So we agreed to trial a new version of the IPS Terminology format:

NOVEMBER 2022:

APRIL 2023:
ANY FEEDBACK FROM USING IT IN PRODUCTION SYSTEMS???
YES!
Previous changes to the file format addressed the issues that they had - so that's good
People were however unhappy that this is being published separately, 
...and via a different mechanism to the usual MLDS distribution method
This creates more work for implementers and NRC's to consume
MLDS is already full of historical non-SNOMED CT content (Resources, etc)
The use-case for Members using the IPS Terminology product (that was originally designed specifically for non-members to use as an intro to SNOMED before getting full SNOMED licence), is that they want to be able to create queries (FHIR value sets, etc) that work for BOTH Members and non Members, allowing Members to transfer data to and from non-Members.  Therefore in order to make this happen, and to be able to test the end to end, they need to be able to test them against not only the FULL SNOMED (that Members are using) but ALSO against the IPS Terminology scope (that non-Members are using).
HAVING DISCUSSED THIS INTERNALLY WE WOULD BE HAPPY TO PUBLISH IPS TERMINOLOGY VIA MLDS (as well as the IPS part of the SNOMED website) - WOULD THIS RESOLVE THIS ISSUE??
 
IN ADDITION, some people are unhappy with the separate IPS URI (eg) "http://snomed.info/ips/999991001000101"
HAVING DISCUSSED THIS INTERNALLY WE WOULD NEED A REALLY STRONG USE CASE TO CHANGE THIS AT THIS POINT - CAN ANYONE PROVIDE ONE, OTHER THAN THAT IT'S A BIT IRRITATING?
.
 
AAT to start publishing IPS Terminology on MLDS AS WELL as on the IPST download site, to allow easier access for Members.

Request was also made for the URI to change from http://snomed.info/ips to something more standard

This was initially rejected internally, as the entire point of this was to distinguish IPST from other SI products
However, Australia then confirmed that Ontoserver CANNOT consume this type of API. 
Peter Williams also suggested that maybe Snowstorm and/or the Browser might not consume it either...
If this is the ALSO the case then we would have a stronger business case to change the URI 
Peter will therefore confirm shortly and we will decide from there...
ONCE ALL DECISIONS MADE WE NEED TO
a)  Inform the community if any changes to be made, and 
b)  Update the SI URI Spec (again if any changes are to be made, or even if we're keeping it as .../IPS/... as this isn't in the spec??

13

SNOMED Release Package causing file path length issues in Windows environments.

 

@Patrick McLaughlin @John Snyder 

WEDNESDAY (MEETING 2) WHEN US NRC IS IN ATTENDANCE:

  • The base file name is 63 characters:  SnomedCT_ManagedServiceUS_PRODUCTION_US1000124_20230301T120000Z.zip

  • When the zip package is decompressed on a windows based computer using 7zip, the base folder name is the same length as the zip package file name:

  • Drilling down into the second level directory, we see that the base folder is duplicated:

  • The user must click through the duplicate folder before actually getting to the Full and Snapshot folders:

    •  

  • The duplicated folder is easily viewable when looking at the file path.

    • C:\Users\snyderjw\Desktop\SNOMED\

    • ...SnomedCT_ManagedServiceUS_PRODUCTION_US1000124_20230301T120000Z\...

    • ...SnomedCT_ManagedServiceUS_PRODUCTION_US1000124_20230301T120000Z

  • To access files in the release package, the SNOMED path and file name can reach up to 218 characters.

    • ..\SnomedCT_ManagedServiceUS_PRODUCTION_US1000124_20230301T120000Z\...

    • ...SnomedCT_ManagedServiceUS_PRODUCTION_US1000124_20230301T120000Z\

    • ...Snapshot\Refset\meta\

Copyright © 2026, SNOMED International