From Months to 48 Hours
Building sustainable administrative data infrastructure for public health surveillance
A state public health agency replaced a slow, person-dependent process for preparing annual hospital discharge data with a governed and maintainable workflow.
New releases became available in 24–48 hours instead of months. The infrastructure supported approximately 70 users across ten programs, and the agency estimated that it recovered its investment within two years.
Most important, automation preserved epidemiologic oversight while allowing specialized staff to spend more time on surveillance, interpretation, and public health action.
-
24–48 hours
From receipt of a new release to availabilityApproximately 70 users
Across the public health agencyTen programs
Using the data for surveillance, evaluation, and reportingApproximately two years
Estimated time required to recover the investment -
The agency received annual Hospital Discharge Data containing hundreds of thousands of patient discharges from acute care hospitals across the state. Records included diagnostic and procedure codes, demographics, dates of care, and charge information.
Programs used the data in different ways:
Injury epidemiologists relied on it for core surveillance.
Occupational health programs used it to verify cases and identify cases missing from other reporting systems.
Community health programs used hospitalization measures for evaluation and grant reporting.
Other programs used selected fields to supplement specialized surveillance systems.
These uses created different requirements for fields and derived variables. In some cases, programs requested different versions of similar information because their surveillance purposes differed.
The records also included patient dates of birth. The agency therefore managed them as a HIPAA limited data set on secure infrastructure with controlled access.
-
Before the project, an epidemiologist prepared each annual release manually. The work could take months and depended heavily on one specialist’s availability and knowledge.
The process reflected a valid concern. Administrative health data cannot be trusted simply because a file loads successfully. Changes in coding, file structure, populations, and source-system practices can introduce errors that require subject-matter expertise to identify.
The transition from ICD-9 to ICD-10 made the incoming files larger and substantially more complex. The agency needed a faster process without losing the epidemiologic judgment that made the data trustworthy.
-
As a member of the project leadership team, I created a working group of epidemiologists representing programs across the agency. These representatives gathered requirements, communicated with program managers and colleagues, and participated in testing.
I had final approval on whether requirements made sense from an epidemiologic perspective. I identified ambiguities, overlapping needs, and requests that required a clearer explanation of their intended use.
When programs requested different transformations of similar information, I worked with their representatives to understand the analytical purpose behind each request. If those conversations did not resolve the issue, I brought the affected programs together to develop a shared solution.
As implementation neared completion, I assumed full ownership of the governance process.
-
The agency’s central data office, public health programs, and IT developed a repeatable production and validation workflow.
Receive
The central data office received each release from the external health information agency and transferred it to IT.
Load
IT ran the automated extraction, transformation, and loading process, generally overnight.
Validate
The team checked whether processing completed successfully and ran additional validation in SAS. Checks included patient counts, admission counts, demographic distributions, and processing errors.
Resolve
The team compared anomalies with the data provider’s release notes. Unresolved issues were escalated to the provider and could result in clarification, updated documentation, or a replacement file.
Release
Once validated, the data became available through the agency’s SQL environment. Users received a notice containing the provider’s release documentation and any additional issues identified by the agency.
The workflow concentrated expert attention on requirements, validation, anomalies, and interpretation rather than repetitive file construction.
-
Program managers approved staff members who needed access. The central data office provided a second review, submitted approved requests to IT, and maintained the authoritative user list.
Each year, managers reviewed the users associated with their programs and removed anyone who had left the agency or no longer needed access.
The team also considered creating a separate database view for every program. We decided against that approach because program requirements changed from year to year. A growing collection of customized views would have become difficult to maintain.
Technical processing, documentation, release communication, and access governance operated as one service.
-
After implementation, new Hospital Discharge Data releases became available within approximately 24–48 hours of receipt.
The infrastructure supported about 70 users across ten programs, including three or four programs that depended on the data for core surveillance.
The benefits included:
Injury surveillance produced final estimates approximately four to six months sooner.
Occupational health staff could verify cases and outcomes and identify additional cases using dependable supplemental data.
Community health programs could use hospitalization measures for evaluation and grant reporting because the data was available on a predictable schedule.
Programs gained access to current information almost as soon as it entered the department.
The agency estimated that the initial investment was recovered within two years. As new annual releases and source changes arrived, the team updated the transformation logic incrementally rather than rebuilding each dataset by hand.
-
Many epidemiologists initially doubted that automation could accommodate the nuances of public health data. Their skepticism reflected previous experiences with technical projects and a legitimate concern that implementation might displace epidemiologic judgment.
The project succeeded because those concerns became design requirements.
Epidemiologists defined what trustworthy data meant. IT staff translated those requirements into a reliable production process. The central data office resolved cross-program questions, coordinated with the source agency, and established governance.
As the agency’s first internally led data automation project of its kind, the work also changed expectations. Other groups began to consider shared, automated infrastructure for their own data resources.
-
Number 1: Treat automation as a capacity investment
Predictable access makes data more useful for surveillance, evaluation, reporting, and time-sensitive decisions. Its value extends beyond staff hours saved.
Number 2: Automate repetition while preserving judgment
Expert review should focus on meaningful anomalies, changing definitions, interpretation, and emerging public health questions.
Number 3: Build requirements around actual uses
Similar requests may support different surveillance purposes. Understanding how a field or transformation will be used is more valuable than collecting specifications alone.
Number 4: Include governance in the operating model
Access approval, annual user review, documentation, release communication, and issue escalation should be designed alongside the technical workflow.
Number 5: Design for change
Source files, coding systems, and program priorities will evolve. Sustainable infrastructure accommodates incremental changes without multiplying custom processes.
-
Public health data modernization works when technical design reflects epidemiologic practice.
A dependable system gives epidemiologists, IT staff, program leaders, and data providers clear responsibilities. It removes repetitive work while preserving the judgment required to produce trustworthy surveillance data.
-
I help state and local public health agencies turn complex administrative health data into reliable, maintainable resources for surveillance and evaluation.
My work connects epidemiologic priorities with government operations, technical implementation, and cross-program governance. The goal is to make current information available sooner while allowing epidemiologists to devote more of their expertise to interpretation and public health action.
If your agency is working through a surveillance data bottleneck, modernization initiative, or cross-program data challenge, I would be glad to learn more.