Medical Code Lookup Tool Evaluation Report v1.0.pdf

PDF 2 MB Posted

Attached to
IDIQ SENTINEL SYSTEM 3. 0: PROGRAM MANAGEMENT ORGANIZATION (PMO) Federal contract opportunity
Solicitation number
RFP75F40124R00181
Issued by
Department of Health and Human Services Food and Drug Administration

About this file

This document is a Government File describing the Medical Code Lookup Tool (CLT) used within the United States Food and Drug Administration's (FDA) Sentinel System to allow investigators to search for and build medical code lists for pharmacoepidemiologic studies. The report evaluates the current functions of CLT and identifies potential improvement opportunities by comparing it to alternative Commercial-Off-The-Shelf (COTS) and open-source code set construction and sharing tools. Key findings include: CLT meets most recommendations from a 2017 academic report on best practices for medical code list methodology, exceeding the other tools analyzed; suggested enhancements to CLT include adding a synonym mapping database, enabling iterative code list refinement, and improving metadata capture. The report concludes that CLT is purpose-built for the FDA's needs and there is currently no open-source or COTS solution that could adequately replace it.

This document also describes a related federal contract opportunity for the Sentinel System 3.0 Program Management Organization (PMO) contract. The PMO contract will provide program/project management and business informatics support services for the broader Sentinel System 3.0 initiative, which will include the Sentinel Coordinating Center, a Sentinel Data Hub, and Federal Partner agreements. The opportunity is being solicited through RFP75F40124R00181 by the Department of Health and Human Services Food and Drug Administration.

View the file

Other files for this federal contract opportunity

Show all 15

On GovTribe

Work with this file on GovTribe

  • Download the original file
  • Contacts named in this file
  • Similar government files
  • Ask GovTribe AI about this file

Text version

The Sentinel System is sponsored by the U.S. Food and Drug Administration (FDA) to p roactive ly monito r the safe ty of FDA-re g ulate d me d ical p rod ucts and comp le me nts o the r e xisting FDA safe ty surve illance cap ab ilitie s. The Se ntine l Syste m is one p ie ce of FDA’s Sentinel Initiative , a long -te rm, multi-face te d e ffort to d e ve lop a national e le ctronic syste m. Se ntine l Collab orators includ e Data and Acad e mic Partne rs that p rovid e acce ss to he althcare d ata and ong oing scie ntific, te chnical, me thod olog ical, and org anizational e xp e rtise . The Se ntine l Coord inating Ce nte r is fund e d b y the FDA throug h the De p artme nt of He alth and Human Se rvice s (HHS) Contract numb e r 75F40119D10037.

Medical Code Lookup Tool Evaluation O p e ra tio na l Ye a r 2024 Enh a nce m e n t

Theresa Clemmons, Product Owner 1

Matti Hautala, MPAff, Senior Technical Project Manager 1

Laura Hou, MS, MPH, Senior Research Analyst1

Jenice Ko, MPH, Research Analyst1

Ashley I. Michnick , PharmD, PhD; Senior Research Scientist1,2

Janavi Patel, MS, Project Coordinator 1

Max Ehrmann, CLT Subject Matter Expert1

Christine Lee Halbig, MPH, Sentinel Infrastructure Program Manager 1

Morgaine Payson, CLT Subject Matter Expert1

1Department of Population Medicine, Harvard Pilgrim Health Care Institute, Boston, MA 2Harvard Medical School, Boston, MA

Version 0.91 March 1 8, 2024 http://www.fda.gov/ http://www.fda.gov/Safety/FDAsSentinelInitiative/default.htm

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt i

Medical Code Lookup Tool Evaluation Operational Year 2024 Enhancement

Table of Contents

TABLE OF CONTENTS ...................................................................................................... I

1 HISTORY OF MODIFICATIONS .............................................................................. I

2 EXECUTIVE SUMMARY

INTRODUCTION

2.1 Historical Context of the Medical Code Lookup Tool

2.2 The Medical Code Lookup Tool’s Re-Evaluation

METHODS

2.3 Scope of the OY2024 CLT Enhancements Project

2.4 Approach to the Re-Evaluation

2.5 Preparations for the Re-Evaluation

KEY CODE SET CONSTRUCTION FUNCTIONALITY: SEARCHING FOR CODES

2.6 Minimum Requirement: Construct an Initial List of Synonyms in Code Searches

2.7 Minimum Requirement: Make Use of Hierarchies in Code Searches

2.8 Minimum Requirement: Iterate on Search Results for Additional Synonyms

2.9 Assessment of Minimum Requirements for Code Searching

KEY CODE SET CONSTRUCTION FUNCTIONALITY: CODE SET RE-USE

2.10 Minimum Requirements: Record and Store Key Metadata

2.11 Minimum Requirements: Facilitate Code Set Multiplication and Re-Use

2.12 Minimum Requirements: Enable Backward-Compatibility and Tracking

2.13 Assessment of Minimum Requirements for Code Set Re-Use

KEY CODE SET CONSTRUCTION FUNCTIONALITY: CODE TOOL USABILITY

2.14 Minimum Requirement: Minimal Setup for Widespread Adoption

2.15 Minimum Requirement: Easily Accessible

2.16 Minimum Requirements: Open-to Source and Public

2.17 Assessment of Minimum Requirements for Tool Usability

RECOMMENDATIONS & CONCLUSIONS

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt i

1 History of Modifications

Version Date Modification Author

0.8 2024-02-16 O rig inal ve rsion (fo r SO C le ad e r re vie w) Se ntine l O p e rations Ce nte r

0.9 2024-03-01 Up d ated ve rsion (fo r FDA initial d raft re view) Se ntine l O p e rations Ce nte r

0.91 2024-03-19 Final d raft (fo r SO C le ad e r revie w) Se ntine l O p e rations Ce nte r

1.0 2024-03-29 Final d raft fo r sub mission to FDA on p ro je ct close Se ntine l O p e rations Ce nte r

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 1

2 Executive Summary The Sentinel Initiative is a crucial part of the United States Food and Drug Administration’s (FDA) national electronic system for monitoring the safety of FDA-regulated medical products, drugs, vaccines, and medical devices through a national distributed data network. The Medical Code Lookup Tool (CLT) is an application within the Sentinel System that was built for the FDA in 2014 to allow investigators to (a) search for diagnosis, procedure, drug, and other medical codes in a standardized manner and (b) build methodologically sound and reproducible code lists for pharmacoepidemiologic insights. Since 2021 alone, CLT has been used to help answer over 150 pharmacoepidemiologic queries posed by the FDA, and it continues to be a crucial component of the Sentinel infrastructure.

To ensure that CLT remains innovative and responsive to the FDA’s needs, we prepared a comparative landscape analysis to summarize the key features and benefits of CLT and contrast them with the following non-CLT commercial-off-the-shelf (COTS) and open-source solutions: CALIBERcodelists & CALIBERlookups; ClinicalCodes.org; Clinical Table Search Service (CTSS); CPRD Code Browser; Data Compass; Find-A-Code; HDR UK Phenotype Library; OHDSI/Athena; OpenCodelists; RxNAV.

While there are no universally accepted standards for medical code lookup and code list creation, CLT meets most recommendations and requirements described in a seminal 2017 academic report by Williams et al. on medical code list methodology1 – meeting more of the recommendations than any of the 10 other solutions analyzed in this report.

Section 8 of this report (Recommendations and Conclusions) gives a detailed, point-by-point assessment of CLT functionality against the Williams et al. report. For any recommendations CLT does not currently meet, there are enhancement descriptions that can be further prioritized with FDA for future development. The highest priority recommendations from the workgroup are listed here:

Minimum Code Set Construction or Sharing Tool Requirement a

Recommended CLT Enhancement

Cod e se t construction too ls should facilitate the initial construction of a list o f synonyms

Ad d ing a synonym map p ing d atab ase to e xp and on the sug g e stion cap ab ility in narrow se arch, to find like cod e s b ased off the o rig inal se arch. This will allow use rs to se e an ad d itional array of cod e s and allow the m to ad d synonyms in anothe r ite ration of the cod e list.

Cod e se t construction too ls should facilitate an ite rative p roce ss b y sug g e sting ad d itional synonyms b ase d on cod e s d iscove re d b y se arching the cod e hie rarchy that are no t found the mse lves d ire ctly

Ad d ing the ab ility to take an orig inal cod e list and ite rate on it b ase d off synonyms or manual re vie w. This will allow use rs to track how a cod e list has e vo lved from the o rig inal se arch via synonyms or custom chang e s b ase d off e xp e rt fe ed b ack. Histo rical cod e use s in ce rtain te rminolog ie s could make id e ntifying cod e re -use e ve n more e fficie nt.

We conclude that CLT is truly purpose-built for the FDA’s needs and there is currently not an open-source or COTS solutions that could adequately replace CLT. While this project did not include full cost comparisons due to the project timeline, available cost estimates are listed in Appendix A. Sentinel Operations Center (SOC) will work with FDA to define future next steps, which could include further cost evaluation, and/or a deeper assessment of the recommended enhancements for their level of effort and impact on workflow.

1 Richard Williams et al., “Clinical Code Set Engineering for Reusing EHR Data for Research: A Review,” Journal of Biomedical Informatics 70 (June 1, 2017): 1–13, https://doi.org/10.1016/j.jbi.2017.04.010.

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 2

Introduction

2.1 Historical Contex t of the Medical Code Lookup Tool

Querying routine healthcare databases such as those in the Sentinel Distributed Database (SDD) requires creating sets of clinical codes that are used in various parts of a pharmacoepidemiologic study, including to identify cohorts, covariates, and outcomes of interest. Often one of the earliest steps in designing studies leveraging secondary data is to construct a clinical code set to help establish a study’s viability and validity. Indeed, previous research has shown that differences in chosen code sets can introduce differences in estimates of disease prevalence by as much as sevenfold.2

Acknowledging the importance of code set creation within the Sentinel System, the United States Food and Drug Administration (FDA) sought a standardized and efficient way to build and manage medical code lists. At the time, the Sentinel System’s Coordinating Center (Harvard Pilgrim Health Care Institute; HPHCI) had a process for creating and using code lists that was manual and labor-intensive as it required tracking unstandardized querying of datasets within a local network folder. Queries would be written in SAS by an analyst and run against the medical terminologies, but to reuse, iterate, or build upon previous searches, the analyst would need to go through SAS files manually to assess the similarities and capacity for re-use and then configure their own query. Additional current features in The Medical Code Lookup Tool (CLT) such as on-screen user tips and comprehensive, easy-to-use hierarchies were also not available or possible within this folder structure.

In 2014, HPHCI evaluated a separate but similar application to CLT that was created by Commonwealth Informatics (CII) for another client, but this application was ultimately deemed to not meet the specific business requirements of HPHCI and FDA, due to its primary focus on searching via hierarchy in lieu of more metadata-centric searches and its limited ability to export customized code lists. Also in 2014, HPHCI conducted additional market research and were unable to identify a Commercial Off the Shelf (COTS) or open-source application that would provide a centralized and standard system for securely tracking and storing diagnosis, procedure, and drug codes.

Without an existing commercial or open-source solution, and without industry consensus as to what an optimal code set construction tool should look like, HPHCI undertook the task of developing a fit-for-purpose tool and in 2014 released the first version of CLT. The continued enhancements and evolution of CLT have provided the Sentinel System with a centralized and standardized tool that delivers the following key functionality through a web-based application:

• Centralized access to a variety of up-to-date and historical medical terminologies that allows users to search specifically and flexibly for diagnosis, procedure, and drug codes in a standardized manner.

2 Sara Muller et al., “An Algorithm to Identify Rheumatoid Arthritis in Primary Care: A Clinical Practice Research Datalink Study,” BMJ Open 5, no. 12 (December 23, 2015):

e009309, https://doi.org/10.1136/bmjopen-2015-009309.

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 3

• Intuitive User Interface and Experience (UI/UX) that does not require knowledge of a programming language, nor installation of any software, while at the same time supporting code set metadata storage.

• Multiple code search and code set storage methodologies, including the ability to export search results and save criteria to enable efficient code set re-use and validation (via a code list library).

During the last 10 years, CLT has been routinely used by research analysts and epidemiologists at the Sentinel Operations Center (SOC) and the FDA to build the initial code lists used to conduct Sentinel queries, and it is an essential component of conducting scientifically rigorous and comprehensive queries.

2.2 The Medical Code Lookup Tool’s Re -Evaluation

Nearly a decade after CLT’s initial development, the FDA contracted with SOC to initiate the 2024 Operational Year (OY2024) CLT Enhancements Project in October 2023. The goal of this project was to evaluate the current functions of CLT and identify potential improvement opportunities available in COTS or open-source code set construction tools and code set sharing technologies. The OY2024 CLT Enhancements Project was a 6-month engagement, culminating in the submission of this final report in March 2024.

Figure 1 below visualizes the timeline of main project activities:

Figure 1. OY2024 CLT Enhancements Project Timeline

As the final deliverable for the OY2024 CLT Enhancements Project, this report has three objectives:

(1) Describe the key functions and features of CLT, Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 4

(2) Compare CLT’s functions and features to alternative COTS or open-source code set construction/sharing solutions, and

(3) Based on the results of the comparative analysis from Objectives 1 and 2 – provide prioritized recommendations to FDA leadership for potential future CLT enhancements based on existing CLT capabilities, alternative functions from external non-CLT solutions, and best practices.

In ensuring CLT remains at the front lines of code set construction, management, and sharing standards, this report will be central to ensuring that the FDA and its Sentinel System remain a paramount resource for rigorous public health surveillance in the United States.

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 5

Method s

2.3 Scope of the OY2024 CLT Enhancements Project

As mentioned above, code set construction, management, validation, and sharing are crucial steps in any rigorous pharmacoepidemiologic study. While this process requires efficient and usable tools, it is also heavily reliant on complex operational processes and procedures involving expert code review and study-specific decisions around code inclusion. This project and report focus on analyzing and describing the functional and technical components (and related process components) of The Medical Code Lookup Tool (CLT) as an application that is central to United States Food and Drug Administration (FDA)’s mission for public health and safety within the Sentinel System.

Accordingly, this report identifies and evaluates potential improvement opportunities for CLT but should not be considered a comprehensive assessment of all available alternatives nor a final recommendation on whether any Commercial Off the Shelf (COTS) or open-source solutions could replace CLT within the Sentinel System. Because there are no universally accepted standards for code set construction tools, and because the Sentinel Operations Center (SOC) team and CLT itself are designed to meet FDA’s specific needs, the comparisons made between CLT and other solutions are primarily focused on CLT uses and functionalities and are not lateral comparisons. It was also outside the scope of this project to obtain licenses and fully trial each of the external, non-CLT solutions. For this reason, the report is focused on potential functional enhancements to CLT and does not address costs associated with either implementing enhancements to CLT or replacing CLT with a COTS or open-source solution.

2.4 Approach to the Re -Evaluation

It has been a decade since the initial design and implementation of CLT and as the technological landscape changes, the SOC has undertaken the necessary due diligence to evaluate potential enhancements to the application and explore alternative solutions. To keep pace with CLT user needs and emerging security vulnerabilities, the SOC CLT System Support Team implements small improvements related to CLT’s Intuitive User Interface and Experience (UI/UX), performance, and security as part of regular software maintenance. With the goal of meeting the evolving needs of CLT users and the FDA, this project provided the opportunity for CLT users and subject matter experts (SME) to evaluate the current performance and utility of CLT compared to alternative COTS and open-source solutions.

The SOC team leading the 2024 Operational Year (OY2024) CLT Enhancements Project objectives (hereafter referred to throughout this report as “we” or “the SOC project team”) chose to utilize a mixed-methods approach, utilizing a brief literature review and focus groups with key stakeholders to establish baselines and frameworks on which to accomplish the objectives of assessing CLT and alternative solution functions, and making recommendations for future enhancements. Key stakeholders included Research Assistants, Research Analysts, Research Scientists, Product Developers, leadership within the SOC, and key FDA leadership assigned to this Project.

After gathering this feedback and using it to frame the next steps, we reviewed CLT’s current documentation and functionalities and performed a landscape analysis of alternative code set construction and sharing tools to meet the Project’s first two objectives. To accomplish the third

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 6 objective, we generated and prioritized recommendations for potential CLT enhancements based on the professional expertise of SOC project team members, the existing backlog of requested developments for CLT, and the FDA’s overarching priority to keep Sentinel as the premier medical product safety surveillance system in the nation.

2.5 Prepara tions for the Re -Evaluation

In order to structure our description of CLT and comparison to potential alternative solutions, we leveraged information gained from interviews and focus groups to classify CLT’s functions into broad categories. To contextualize our findings on functions and features, we referenced a list of “minimum requirements” for code set construction and sharing tools. Finally, we used subject matter expertise to identify a select list of exemplar COTS and open-source alternative solutions to which we could compare CLT’s function and features. Below, we describe our process for classifying tool functionality, establishing minimum requirements, and choosing alternative solutions for comparison.

2.5.1 Classifying Tool Functionality

Our interviews and focus groups revealed that Sentinel’s CLT is a rich resource for constructing, managing, and sharing clinical code sets. These viewpoints helped us to establish functions within CLT that were key to users’ abilities to accomplish their tasks. We were able to formulate a list of CLT’s benefits and potential areas of opportunity from key stakeholders, as shown below in Figure 2.

Figure 2. Summary of Focus Group Findings on Current Benefits and Areas of Opportunity for CLT

From these areas of benefit and opportunity, we derived three key functions around which to frame our description of CLT and comparison to alternatives:

(1) Code search, to encompass the benefits and opportunities of: basic and advanced search flexibility, search result summaries, and more advanced search features

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 7

(2) Code set re-use, to encompass the benefits and opportunities of: batch search features, code list library storage and access, and code set comparison features

(3) Tool usability, to encompass the benefits and opportunities of: customizability, context-specific help menus and tooltips, speed/performance, documentation, and integration

For each functional area, this report details specific features, highlighting similarities and differences to alternative code set creation and code set sharing tools available on the market.

2.5.2 Establishing a Baseline of Minimum Requirements

As mentioned earlier, code set construction is a crucial component of any rigorous pharmacoepidemiologic study, and poor practices can lead to large impacts on the ultimate accuracy and precision of findings. Thus, in describing functions of CLT and its alternatives, we sought industry guidance to use as a baseline by which we could establish minimum requirements. Despite its clear importance, guidance on the creation and maintenance of clinical code sets (and even a base recognition of its importance) is limited.

Considering this clear gap in the literature, a team of researchers from the University of Manchester took on the task of developing recommendations for best practices in clinical code set management – both in practice and in software implementation. In their seminal 2017 methodological article “Clinical Code Set Engineering for Reusing Electronic Healthcare Record (EHR) Data for Research: A Review”, Williams and colleagues perform a review of the literature and provide 13 recommendations by which code set construction tools and code set sharing platforms should abide.3 Neither the authors of the Williams et. al review nor the SOC Project team found any such comprehensive review or set of recommendations prior to 2017, and since its publication the Williams et. al. review has been cited nearly 40 times.4

Indeed, our continued search for a baseline on which to judge CLT and alternative code lookup solutions revealed that even though the “Guidelines for Good Pharmacoepidemiology Practices” report by the International Society of Pharmacoepidemiology (ISPE) notes that “clear operational definitions of disease state, exposures, health outcomes, and other measured risk factors for outcome” should be included in the methods section of protocols, it fails to provide further details regarding suggested best practices or preferred methods for creating such definitions.5 Other reporting guidelines specific to pharmacoepidemiology that have been published since the Williams et. al. report, including RECORD- PE6 and STaRT-RWE7, specify that studies should clearly explain how codes and algorithms used in the study were derived, but similar to the ISPE report these guidelines do not provide detail on how those code sets or algorithms should be generated.

With these limitations in mind, we have chosen to use the Williams review to structure this report’s assessment of CLT’s current and future state functionality and use. Hereafter, we will refer to the Williams et. al report’s 13 recommendations for code set construction tools and code set sharing

3 Williams et al., “Clinical Code Set Engineering for Reusing EHR Data for Research.”

4 Data from 28 February 2024 Web of Science™ Clarivate™ © Clarivate 2024. All rights reserved.

5 Public Policy Committee, International Society of Pharmacoepidemiology, “Guidelines for Good Pharmacoepidemiology Practice (GPP),” Pharmacoepidemiology and Drug Safety 25, no. 1 (January 2016): 2–10, https://doi.org/10.1002/pds.3891.

6 Sinéad M. Langan et al., “The Reporting of Studies Conducted Using Observational Routinely Collected Health Data Statement for Pharmacoepidemiology (RECORD-PE),” BMJ 363 (November 14, 2018): k3532, https://doi.org/10.1136/bmj.k3532.

7 Shirley V. Wang et al., “STaRT-RWE: Structured Template for Planning and Reporting on the Implementation of Evidence Studies,” BMJ 372 (January 12, 2021): m4856, https://doi.org/10.1136/bmj.m4856.

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 8 platforms as “Minimum Requirements”, and the report will detail the functions and features of CLT and non-CLT solutions with respect to whether they meet each of the 13 recommendations.

Table 1 below contains the 13 minimum requirements in the far-left column, followed by the category of CLT functions to which that minimum requirement applies.

Table 1. Minimum Requirements for Code Set Construction Tools and Applicable CLT Functions

Minimum Requirement s Applicable CLT Function

Code set construction tools should be open source, publicly available, and easily accessible Tool Usability Code set construction tools should be backward compatible with code sets produced with earlier versions of the phenotyping tool

Code Set Re-Use

Code set construction tools should have minimal setup to facilitate widespread adoption Tool Usability Code set constru ction tools should make use of the hierarchy of clinical dictionaries to assist in the searching and selection of codes

Code Search

Code set construction tools should facilitate the initial construction of a list of synonyms Code Search Code set constru ction tools should facilitate an iterative process by suggesting additional synonyms based on codes discovered by searching the code hierarchy that are not found themselves directly

Code Search

Code set construction tools should facilitate the various stages of the review process from the selection of the initial list of synonyms, to the review of the included and excluded codes

Tool Usability

Code set construction tools should make [the code set construction and review] process as simple and q uick as possible

Tool Usability

Code set construction tools should facilitate and encourage the creation of multiple sets of codes for sensitivity analyses

Code Set Re-Use

Code set construction tools should record metadata such as: the initial list of sy nonyms, the excluded codes, the purpose for the set and the author

Code Set Re-Use

Code set construction tools should facilitate the reuse, validation and sharing of codes sets, not simply their construction

Code Set Re-Use

Code set sharing platforms should be discoverable, maintained indefinitely and support versioning

Code Set Re-Use

Code set sharing platforms should support the storage of metadata alongside the code set Code Set Re-Use a Minimum requirements taken from recommendations in the 2017 Williams et. al. review of code set creation and management

2.5.3 Choosing Alternative Solutions for Comparison

Once we had classified CLT functions into three minimum categories and established a framework of minimum tool requirements on which to base our comparisons, we needed to compile the set of alternative COTS and open-source code set construction, management, and sharing solutions for comparison to CLT’s functions and features. We leveraged subject matter expert (SME)s and our non-systematic literature review to inform this compilation because no comprehensive assessment of all available tools existed, and because our literature review confirmed that code set construction tools and code set sharing platforms varied greatly by institution. While there are many tools (either COTS or open-source) available with a wide span of purposes (e.g., medical billing and treatment), to ensure the most analogous comparisons, we chose to restrict our focus to tools primarily used in secondary data evaluation, assessment, and/or research.

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 9

As shown in Figure 3. we identified 10 alternative COTS or open-source code set construction and code set sharing tools that could serve as potential alternatives to the Sentinel System’s CLT. Each of the solutions identified for this Project presented a unique set of functions and features that set them apart from one another. This analysis focused on comparing code list tools themselves and, unlike the Williams report, does not include an analysis of highly customized or manual code list systems that would not be scalable to the Sentinel System.

Figure 3. COTS and Alternative Code Set Construction and/or Sharing Tools

After extensive review, the number of alternative solutions assessed in detail for this project was reduced to include products which offered functions and features equivalent or better than those currently implemented within CLT, so that if a recommendation were made to acquire or implement the alternative solution, it would not depreciate any of the value or efficiency in the Sentinel System’s current code set process (CLT). Ultimately, four COTS or alternative solutions were used as primary comparators throughout the assessment: Athena, Find-A-Code, OpenCodeList, and CALIBERcodelists.

A summary containing links to these products and information on funding, development, and location of use is found in Table 2 below.

Table 2. Summary of Alternative Solutions Compared to CLT in This Report

Product Funder Developer Location of Use

Athena Public-private collaboration

Observational Health Data Sciences and Informatics (OHDSI) International

CALIBERcodelists Wellcome Trust + NIHR Farr Institute International Find-a-Code innoviHealth innoviHealth International

OpenCodelists OpenSAFELY University of Oxford for the Bennett Institute for Applied Data Science International

Summary tables for each COTS and open-source solution were created to capture the breadth of our comparative landscape analysis of code set construction tools and code set sharing platforms. These https://athena.ohdsi.org/search-terms/start https://caliberanalysis.r-forge.r-project.org/ https://www.findacode.com/ https://www.opencodelists.org/

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 10 summary tables can be found in Appendix A and describe the 10 tools (plus CLT) noted above in Figure 3 in greater detail, highlighting key pieces of high-level product information including a detailed list of which medical terminologies are included in each of the solutions as implemented. Of note, the OY2024 CLT Enhancements Project and this report do not focus on analyzing or enhancing the medical terminologies themselves, but rather the tools that organize and distribute them. Because terminology scope, cost, and other details are important factors when comparing tools, we have opted to include this additional information in Appendix A.

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 11

Key Code Set Construction Functionality : Searching for Codes As shown in Table 1 in the Methods section above, three of the 13 minimum requirements for code set construction tools or sharing platforms apply to this function and more than one-third of the benefits and opportunities identified by our focus groups (summarized in Figure 2 above) pertained to code searches.

Medical codes are also referred to as medical terms, and medical terms are found in different medical terminologies, which are languages that were created to describe medical concepts. For example, The Medical Code Lookup Tool (CLT) (as of version 13.6.3; released March 2023) allows users to perform searches for medical terms in 11 different medical terminologies, as shown in Table 3 below.

Table 3. Summary of Medical Terminologies Available in CLT

Code Type Developer Vendor Terminology Update Frequency a Drug USFDA First DataBank NDCb Quarterly Drug NIH Developer RxNormc None scheduled Diagnosis CDC Optum ICD-9-CMd None scheduled Diagnosis CDC Optum ICD-10-CMd Quarterly Procedure CMS Optum HCPCS Levels l and ll d Quarterly Procedure CDC Optum ICD-9-CMd None scheduled Procedure CMS Optum ICD-10-PCSd Quarterly Clinical SNOMED International IHTSDO SNOMED CTd None scheduled Clinical Regenstrief Institute Developer LOINCc Semi-annual Clinical CDC Developer Codabar c Semi-annual Clinical ICCBBA Developer ISBTc Semi-annual

CDC: Centers for Disease Control and Prevention; CM: Clinical Modification; CMS: Centers for Medicare and Medicaid Services;

HCPCS: Health Care Procedure Coding System (includes Current Procedural Terminology [ CPT] as Level I); ICCBBA: International Council for Commonality in Blood Banking Automation ; ICD: International Classification of Diseases; IHTSDO: International Health Terminology Standards Development Organisation ; ISBT: Information Standard for Blood and Transplant ; LOINC: Logical Observation Identifiers Names and Codes; NIH: National Institutes of Health; NDC: National Drug Code; PCS: Procedure Coding System; SNOMED CT: Systematized Nomenclature of Medicine Clinical Terms; USFDA: U.S. Food and Drug Administration a Frequency of terminology updates is determined in part by FDA and SOC priority and needs b Terminology not inherently hierarchical; CLT uses vendor -provided hierarchy c Terminology not inherently hierarchical d Terminology inherently hierarchical

These medical terminologies are used to classify and organize medical data using standardized terms, codes, and/or definitions. Each terminology has its own refresh schedule based on its source (see Table 3 for refresh schedules), and users select which version of a terminology they would like to use as a first step in every search.

Each terminology is set up in a database table that CLT uses as the basis for each search (see an example for the National Drug Code (NDC) terminology in Figure 4 below and an example for the International Classification of Diseases, 10th Revision, Clinical Modification (ICD-10-CM) terminology in Figure 5 below). Medical terminologies generally have several variables (represented as columns in the database table) associated with each term. (The variables available in a terminology are highly dependent and reflective of the terminology itself and are not the focus of this report. For more

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 12 information on the variables available in each terminology implemented in CLT, see CLT documentation.)

CLT can search not only the terms in a database table that are contained in rows, but also in each term’s variables. Variables associated with each term are organized in columns of the database table (as shown in Figure 4 and Figure 5 below) and include attributes of a term such as the brand and generic name in the NDC terminology (or Full Description and Levels 1-4 Terms for ICD-10-CM diagnosis codes). For terminologies that have a hierarchical structure, the database table also stores each level of the hierarchy as a variable, thus allowing hierarchies to be leveraged when performing searches. Table 3 details which terminologies in the CLT are inherently hierarchical, not hierarchical but have third-party hierarchies imposed upon them, or have no hierarchy available. As discussed below, hierarchies can aid in code searches by not only providing more variables on which to search, but also by allowing more advanced search techniques (see Section 2.7 for more details on how hierarchies can improve code searches).

Figure 4. Example Database Table Excerpt for NDCs

Figure 5. Example Database Table Excerpt for ICD-10-CM Diagnosis Codes

Once a user identifies the medical terminology and version in which they’d like to search, they can apply their specific search criteria and create what is known as a “set.” If a user needs to combine one or more sets that use the same terminology, they can do so by creating a superset. Supersets can then be further combined into requests, which can be exported and shared as needed. Within the Sentinel System, requests are often shared with the FDA query request members for further validation and curation as part of the SOC’s operational processes for code set management.

In this section, we describe features within CLT that perform the code search function and compare its performance with respect to the relevant minimum requirements: (a) constructing an initial code set,

(b) making use of terminology-specific hierarchies, and (c) facilitating code set iteration from search results.

2.6 Minimum Requirement : Construct an Initial List of Synonyms in Code Searches

Because the primary purpose-built function of CLT is to construct initial lists of synonyms, CLT excels in this capacity. Below, we discuss unique features that allow users to specify the scope of their search (known as performing broad, narrow, or code searches), additional features that allow implementation

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 13 of more complex search criteria (use of Boolean logic and wildcards), and features that improve search efficiency (such as the auto-complete, suggestion, and batch search dialogs).

2.6.1 Initial Search in CLT: Scope

Once the appropriate terminology and version are selected, users choose whether to perform a primary (broad), supplemental (narrow), or a specific code search (see Figure 6 below).

Figure 6. Search Features in CLT

The names of these three features correspond to the level of information within a terminology that is searched. In broad searches, users select a scope of variables within a terminology to which search criteria are applied, which can be as broad as all the variables for that terminology to as narrow as one variable in the terminology (see Figure 7 below). In narrow searches, users may also choose to leverage variables in a terminology, but instead of looking for any matches across several variables, they specify exact strings in specific variables and choose whether to include or exclude those strings from search results (see Figure 7 below). As suggested by its name, the code search feature allows users to search for a specific code, as opposed to searching for any other variable in a terminology.

Figure 7. Selecting Scope in Broad and Narrow Searches

Being able to specify the search scope is crucial when constructing initial lists of synonyms for Sentinel projects, since projects may not have any pre-conceived notion of a clinical concept prior to constructing an initial list of synonyms. An example of the necessity of specifying search scope is

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 14 highlighted in Figure 4 above. Considering this figure as an excerpt of search results after a broad search for “methotrexate,” (instead of as an excerpt of an underlying database table as above), one will notice that the first result is not in fact the drug methotrexate, but rather an antidote for this drug.

Oftentimes, a code for an antidote may be sufficient evidence of having received a drug of interest, and it may be included in a code list. The only reason this drug was included in CLT’s search results for “methotrexate” is because the user chose to include the third-party hierarchical terms in their search, which included a note that levoleucovorin is used as a methotrexate antidote. This important feature is unique to CLT from other products reviewed and provides users with a more comprehensive and rigorous code set.

2.6.2 Initial Search in CLT: Complex ity

All three search features allow the user to make search criteria more complex by joining multiple text strings with logical operators (& [“AND”] and | [“OR”]), specifying the position in the variable to which the search criteria should be found, and indicating whether the search criteria should be case-sensitive.

Furthermore, broad searches in CLT allow users to select at what position in the variable they would like their search criteria to appear.

Furthermore, while other tools leverage the use of quotes to indicate whether search results should be exact or approximate matches, CLT allows much more specificity by allowing the user to pre-specify whether searches should be exact or approximate within the search dialog. In fact, CLT even cleans search criteria when performing approximate (or “fuzzy”) searches by stripping search criteria of parentheses, periods, forward slashes, and various other non-search criteria before returning results.

Ignoring extraneous characters and white space in search criteria is an efficient and useful way for users to quickly implement complex search criteria.

2.6.3 Initial Search in CLT: Efficiency

Another unique feature of CLT are the methods in which it allows one to populate search criteria. All three features allow users to push values from a prior search’s result set into another search (see the “Minimum Requirement: Iterate on Search Results for Additional Synonyms” section below for more details). In addition, users performing narrow searches can leverage CLT’s “auto-complete” feature to populate search criteria.

Alternatively, the suggestion, search, and batch search dialogs within the narrow search feature offer even more ways of populating search criteria. Specifically, the batch search dialog has no comparison among other tools on the market. This feature allows users to enter numerous search strings separated by a new line and then choose whether to accept or reject the results to be used as input for a new narrow search.

A final unique feature to CLT is its ability to remove formatting and validate criteria when performing batch or code searches. These features (not available in other tools reviewed) were designed with large and complex queries in mind and allow users to input long search criteria with minimal effort.

2.6.4 Initial Search: Comparison to Other Tools

Several other code list creation tools exist on the market that claim to construct initial lists of synonyms just like CLT. However, many of the unique features in CLT related to search specificity and efficiency are due to its built-for-purpose nature for FDA public health surveillance, and thus are not available in these other tools. Nevertheless, there are some distinct features available in potential competitor tools

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 15 that CLT may benefit from implementing. Below, we briefly detail some of these features and our assessment of whether they might benefit CLT.

No available tools allow users to choose the scope of their search before performing it, though many do allow users to use quotes to specify whether their search results should be exact or approximate matches. For example, while OpenCodelists’ text string search is similar to CLT’s broad search in that it searches across all the available information for a given term (even if that information is less than what is available in CLT), OpenCodelists.org does not have a feature similar to CLT’s narrow search, which allows more specific searching of certain variables in a terminology. However, it does offer a feature similar to CLT’s “code search,” which can be accomplished by including the prefix “code:” to a text string search.

As an example, the ATHENA tool from Observational Health Data Sciences and Informatics (OHDSI) does allow users to specify search scope, but the scope is not unique to the selected terminology, since the scope variables leverage OMOP’s framework of limiting search results to certain domains, concepts, classes, vocabularies, or pre-specified validities. The Athena search tool comes the closest to CLT in its ability to limit the scope in which to apply search criteria, but in this open-source tool, the variables are constant across all terminologies [as the variable come from the Observational Medical Outcomes Partnership (OMOP) framework], thus making them less specifically applicable to a given terminology.

The benefit to using this more general framework is that it allows searches across terminologies, which CLT cannot perform.

Like CLT, other tools available on the market also allow the use of logical and/or Boolean operators and wildcards. Others go a step further and allow even more complex searching by leveraging “Regular Expressions” (one such tool is the CALIBERcodelists R package). In this area, CLT may benefit from enhancements to allow even more complex search criteria. Currently, CLT can handle only a limited range of wildcards and only in certain terminologies. Other tools allow more wide usage of wildcards, such as the CALIBERcodelists R package, which allows users to leverage the full scope of wildcards given its reliance on the R software. While this is a potential benefit in allowing more complex search criteria, its relative inaccessibility because of its reliance on a specific software makes it less desirable.

A final unique feature present in tools like the UK HDR Phenotype Library is the ability to integrate work done in other areas into search criteria. In this competitor tool, users can specify as an inclusion criterion that a result set have one or more peer-reviewed outputs associated with it. While this is a unique feature not available in CLT, it also means that the search result set is a pre-curated list, meaning that users may miss the inclusion of valuable codes. CLT’s focus on raw and un-curated search results ensures that projects utilize the full range of available codes. Integrating the ability to reference outside and potentially peer-reviewed code sets and/or phenotypes might be considered a valuable addition to a code set validation tool like the United Kingdom (UK) Health Data Research (HDR) Phenotype Library but is not applicable for a tool like CLT that is primarily focused on code set creation.

A potential area of improvement would be the ability for users to search previously created sets within CLT and return them as part of search results.

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 16

2.7 Minimum Requirement : Make Use of Hierarchies in Code Searches

2.7.1 Hierarchies in CLT

In CLT, the hierarchical nature of certain terminologies can be leveraged in two ways: 1) using text strings to look in the names of hierarchies for matches, and 2) using the “hierarchical view” feature to browse hierarchies. For the former, hierarchies stored as variables in a terminology’s database table can be searched with broad or narrow searches. This feature is shown in Figure 7, where an example of a broad search allowing results matching search criteria in third-party NDC hierarchies [produced by First Databank and known as Enhanced Therapeutic Classification (ETC) terms] can be included. For the latter, users can leverage a visual depiction of a term’s place in a hierarchy to a) explore the hierarchical structure, b) confirm that a search criterion is resulting in the desired places in the hierarchy, or c) specify search criteria. When the hierarchical view is used as described in (c), it is considered a special type of narrow search. Note that while some terminologies are natively hierarchical [e.g., ICD-10 or Healthcare Common Procedure Coding System (HCPCS)], third party organizations have produced hierarchical organization for others (as in the First Databank-produced ETC for NDCs).

These third-party hierarchies aid in initial and iterative code searches and are a valuable addition to such terminologies. Table 3 details the hierarchical nature of each terminology available in CLT.

Figure 8 is an example of what the hierarchical view can look like in CLT when applied to the NDC terminology.

Figure 8. Hierarchical View Excerpt in the NDC Terminology

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 17

Exploring the hierarchical structure for terminologies that have a hierarchy can be a good way to include additional terms in the initial list of synonyms. Without this third-party source, users would not be able to explore the NDC terminology in an organized format that would allow them to include terms they may not have considered including in an initial list of synonyms. An example of an instance in which this can be helpful is when a Sentinel project specifies a class of drugs for which to make a code list but does not specify a specific drug. Being able to browse these third-party hierarchies can make returning initial lists extremely efficient and all-inclusive, especially when new drugs are being added to drug classes frequently. CLT’s unique ability to use these hierarchical node selections as input search criteria is another example of how this feature can accomplish the “iterate on search results” minimum requirement function (see section 2.8 for more details on this minimum requirement).

2.7.2 Comparison to Hierarchies in Other Tools

Some tools (such as Find-A-Code) also allow users to browse code hierarchies after landing upon a result, but unlike CLT they cannot use a hierarchical view to update search criteria.

OpenCodelists (from the OpenSAFELY initiative) allows users to browse hierarchies (when applicable to the terminology) via their “Tree Tab” feature. As in CLT’s “hierarchical view” feature, this OpenCodelists feature shows all of the codes in a set in the context of other codes in the coding system, which can be helpful for seeing whether there are any accidental gaps or additions in the set.

Some available tools on the market are able to leverage hierarchies in terminologies for which CLT has not yet implemented hierarchical searching, but CLT is continually updated and maintained in coordination with the FDA Sentinel Core Team and can add these features (as was done recently for the CPT and HCPCS terminologies). Note that HCPCS Level I is synonymous with Current Procedural Terminology (CPT) codes.

Other tools include terminologies that are inherently hierarchical but that CLT has not implemented.

For instance, because the Sentinel System was initially designed to query primarily administrative claims data, it uses the NDC terminology as its main drug code, since claims in the US are billed using NDCs. However, other data sources such as EHR may use drug terminologies that are inherently hierarchical and thus don’t require the purchase of a third-party hierarchy, such as the Anatomical Therapeutic Chemical (ATC) terminology. As the Sentinel System expands its use of non-claims data sources, including the ATC terminology (and potentially a cross-walk between ATCs and NDCs) is an area for potential future enhancement in CLT that would make it an even more comprehensive tool in its ability to leverage hierarchies in code list creation. Despite its potential for future expansion of hierarchical structures, CLT’s features when taken as a whole make it an effective, purpose-fit tool for Sentinel’s code list creation process.

2.8 Minimum Requirement : Iter ate on Search Results for Additional Synonyms

2.8.1 Search Iteration in CLT: Pushing from Search Results

One crucial feature when considering the ability to iterate on initial search results for finding additional synonyms in a code list is CLT’s “push from results” feature, which can be applied in both broad and narrow searches. With this function, users can input de novo text strings in a broad search, then use the results from that search as input criteria for a new broad or narrow search.

Medical Code Lookup Tool Evaluation | O p e rational Ye ar 2024 Enhance me nt 18

An example of when this is useful is when using approximate (or “fuzzy”) search criteria.

This is the start of the file's text. The full file is on GovTribe.

File details come from the government source that posted it. Updated .