Objective and Study Design
Objective: to quantify real-world agreement in pure-tone average (PTA) measurements between Soundtrace audiograms and employees' historical occupational audiograms obtained within a two-year comparison window.
This is a retrospective field agreement study. Each employee's Soundtrace PTA was paired with the PTA from their most recent historical occupational audiogram completed by an external provider within the prior two years. Historical tests were performed using a wide range of equipment, booths, mobile vans, and technician practices, and are not treated as a gold standard: they are the real-world baseline an employer inherits. Accordingly, this analysis measures agreement between the two measurement processes, not absolute audiometric accuracy or equivalence to a reference audiometer.
The population spans 9,556 employees, 9,496 eligible matched comparisons, more than 250 employers, more than 300 calibrated Soundtrace audiometers, and more than 275 Soundtrace certified facilitators trained under the Soundtrace physician and audiology training model, including CAOHC certified facilitators in Oregon, Washington, and other jurisdictions where certification is required.
PTA Definition. For this analysis, PTA is defined as the arithmetic mean of hearing thresholds at 2,000, 3,000, and 4,000 Hz for each ear. This is the same three-frequency average used in OSHA's Standard Threshold Shift (STS) calculation under 29 CFR 1910.95. The present analysis compares the absolute 2,000/3,000/4,000-Hz averages between Soundtrace and historical audiograms; it does not independently determine STS status or evaluate all requirements involved in an OSHA STS determination.
The primary analysis uses all eligible matched comparisons with no statistical outlier removal. No observation was excluded based on the magnitude of disagreement. Percentile-based outlier exclusion appears only in a clearly labeled sensitivity analysis.
Primary Results: All Eligible Comparisons, No Outlier Removal
Across all 9,496 matched comparisons, the mean paired difference (Soundtrace PTA minus historical PTA) was -2.68 dB, with a standard deviation of 7.37 dB and a 95% confidence interval for the mean of -2.83 to -2.53 dB. A one-sample t-test of the mean difference against zero yields p < 0.001. Agreement rates: 46.9% of comparisons within ±3 dB, 70.7% within ±5 dB, and 91.4% within ±10 dB.
The mean difference of -2.68 dB is smaller than a single 5 dB audiometric threshold step and substantially below the 10 dB change in the average thresholds at 2,000, 3,000, and 4,000 Hz used by OSHA as the threshold for an STS. This comparison provides regulatory context for the magnitude of the observed average difference; the study does not itself assess STS classification.
In Bland-Altman terms, the mean bias is -2.68 dB with 95% limits of agreement of approximately -17.1 to +11.8 dB, computed as the mean difference plus or minus 1.96 standard deviations of the untrimmed paired differences. The wide limits reflect the untrimmed tails of the distribution, examined in detail below, rather than the behavior of the typical comparison: the central 91.4% of all comparisons fall within ±10 dB.
Distribution of PTA Differences (Untrimmed Primary Population)
Soundtrace minus historical PTA (dB). All 9,496 eligible matched comparisons within a 2-year window, no outlier removal
-2.68 dB
Mean difference
7.37 dB
Standard deviation
< 0.001
p-value
91.4%
Within ±10 dB
Untrimmed primary population. Bars show a normal distribution fitted to the observed summary statistics (mean -2.68 dB, SD 7.37 dB, n = 9,496); the observed distribution is more sharply peaked with a small number of extreme values extending beyond the plotted range. Agreement percentages are computed from the observed data. Within ±3 dB: 46.9%.
Observation Flow
The path from the original employee population to the analyzed comparisons is shown below. Eligibility required a matched pair: a Soundtrace PTA and a historical occupational audiogram PTA for the same employee within the two-year window. Exclusions at this stage reflect missing or unmatchable data, not disagreement magnitude. The 8,069 and 6,855 populations below are sensitivity-analysis populations, not exclusions from the primary study population.
| Population | N | Basis |
|---|---|---|
| Employees analyzed | 9,556 | Employees with a Soundtrace audiogram and at least one historical audiogram |
| Eligible matched comparisons (primary) | 9,496 | Matched PTA pairs within the 2-year window; no exclusions for disagreement magnitude |
| 5th-percentile outlier exclusion (sensitivity-analysis population) | 8,069 | Predefined 5th-percentile outlier exclusion criterion; robustness check only |
| 10th-percentile outlier exclusion (sensitivity-analysis population) | 6,855 | Predefined 10th-percentile outlier exclusion criterion; robustness check only |
Magnitude of Disagreement
Rather than removing extreme disagreements, the primary analysis quantifies them. The distribution of absolute PTA differences across all 9,496 untrimmed comparisons:
| Absolute PTA Difference | Share of Comparisons | Approximate N |
|---|---|---|
| Within ±5 dB | 70.7% | 6,714 |
| > 5 to 10 dB | 20.7% | 1,966 |
| > 10 dB | 8.6% | 816 |
Sensitivity Analysis: Percentile-Based Outlier Exclusion
The primary analysis included all 9,496 eligible matched comparisons without outlier removal. To assess the influence of extreme discordant observations, secondary sensitivity analyses were conducted using predefined 5th- and 10th-percentile outlier exclusion criteria. These analyses were performed only as robustness checks and do not replace the untrimmed primary analysis. The number of observations remaining under each criterion is shown below.
The estimated mean difference remained relatively stable across the untrimmed and percentile-restricted analyses, ranging from -2.68 dB to -2.16 dB, while dispersion decreased as increasingly extreme observations were excluded. This sensitivity analysis indicates that the estimated average offset is not driven primarily by the extreme tails of the distribution.
| Metric | Untrimmed Primary | 5th-Percentile Outlier Exclusion | 10th-Percentile Outlier Exclusion |
|---|---|---|---|
| Comparisons (N) | 9,496 | 8,069 | 6,855 |
| Mean Difference | -2.68 dB | -2.25 dB | -2.16 dB |
| Standard Deviation | 7.37 dB | 3.70 dB | 2.94 dB |
| Within ±3 dB | 46.9% | 52.2% | 58.7% |
| Within ±5 dB | 70.7% | 78.6% | 82.6% |
| Within ±10 dB | 91.4% | 97.6% | 100.0% |
| p-value (t-test vs 0) | < 0.001 | < 0.001 | < 0.001 |
Per-Ear Results (Untrimmed Primary Population)
Each ear was analyzed independently across the untrimmed population. The two ears show closely matching offsets, dispersion, and agreement rates, an internal replication of the combined result: two separate measurement series, analyzed separately, produce the same small offset and the same spread.
| Metric | Left Ear | Right Ear |
|---|---|---|
| Mean Difference | -2.39 dB | -2.96 dB |
| Standard Deviation | 8.74 dB | 8.46 dB |
| p-value | < 0.001 | < 0.001 |
| Observed Range | -95.00 to +100.00 dB | -98.33 to +95.00 dB |
| Within ±3 dB | 38.9% | 38.5% |
| Within ±5 dB | 70.6% | 70.1% |
| Within ±10 dB | 89.7% | 90.0% |
Sources of Variability
Differences between Soundtrace PTAs and historical PTAs can arise from several sources: different equipment types and headphone models across providers, different testing environments (sound booths, vans, offices), variability in technician training and instructions, subject factors such as attention, fatigue, or temporary hearing fluctuation, the method by which the historical audiogram was entered or imported into the record, and the time interval between the two tests, up to two years in this window.
The per-ear observed ranges extend to differences of 95 dB or more. Differences of that magnitude are larger than most clinically plausible changes in hearing over a two-year interval, which suggests contributions from record-level factors such as data entry or import of historical audiograms. The per-comparison causes were not adjudicated in this analysis, and no observation was removed from the primary result on that basis; the sensitivity analysis quantifies their aggregate influence instead.
Comparison With Published Audiometric Datasets
The most direct external reference points are published studies of pure-tone threshold agreement: test-retest studies in controlled clinical settings, validations of automated audiometry against manual audiometry, and industrial audiometry reliability studies. A systematic review and meta-analysis of automated threshold audiometry reported average absolute differences of approximately 2.9 dB between automated and manual thresholds, with test-retest reliability comparable to manual audiometry [1]. Clinical test-retest studies report standard deviations of roughly 3 to 6 dB even when the same subject is retested on the same equipment under controlled conditions [2]. Foundational work on industrial audiometry documented substantially larger variability under field conditions, with technician practice, equipment, and test environment identified as major contributors [3]. Validation studies of automated audiometry outside sound-treated environments report mean differences within a few dB of zero and correspondence of roughly 90 to 98% within ±10 dB [4, 5].
The published figures come from controlled or semi-controlled comparisons, typically same equipment, short retest intervals, and curated data. The present study is a harder comparison: uncurated field data, baselines from hundreds of different providers on unknown equipment, and intervals up to two years. Read against that difference in difficulty, the results are consistent with the literature. The untrimmed primary result, 91.4% within ±10 dB with a mean bias of -2.68 dB, is broadly consistent with the roughly 90 to 98% correspondence reported in controlled boothless validation studies [4, 5]. And the sensitivity analyses show what the comparison looks like as the uncontrolled tail is removed: under the 10th-percentile outlier exclusion criterion the standard deviation falls to 2.94 dB, in the range of the 3 to 6 dB reported for same-equipment clinical retests [1, 2]. These published figures come from studies with differing populations, protocols, and agreement metrics, so the comparisons provide context for interpreting the present results rather than direct equivalence claims.
A plausible contributor to this stability is the standardization applied to every Soundtrace test in the dataset: audiometers calibrated and tracked by serial number with the calibration record bound to each audiogram, per-test ambient noise monitoring that enforces permissible noise levels at the moment of testing, facilitators certified through the Soundtrace physician and audiology training model with CAOHC certification where jurisdictions require it, and physician and audiologist review of results. Each of these controls addresses a variability source identified in the industrial audiometry literature [3].
"We publish the untrimmed numbers first because that is what honest field data looks like. Across nearly ten thousand uncurated comparisons, nine in ten Soundtrace audiograms landed within 10 dB of a prior provider's result, and the mean difference stayed under 3 dB under every percentile-based outlier criterion. This study does not isolate the effect of any single control, but that consistency is what we designed our calibration management, ambient noise monitoring, and facilitator training program to support." Dr. Subinoy Das, Chief Medical Officer, Soundtrace.
Limitations
The analysis used the 2,000/3,000/4,000-Hz PTA but did not include the individual frequency-specific thresholds underlying that average. Accordingly, frequency-specific agreement could not be evaluated. Although this PTA uses the same frequencies incorporated into OSHA's STS calculation, this study did not evaluate employee-level STS classification concordance between Soundtrace and historical providers. Median paired differences and a per-comparison Bland-Altman plot require the underlying paired observations and are not reported here; the Bland-Altman mean bias and limits of agreement are computed from the untrimmed summary statistics. Associations between disagreement and time between tests, device, facilitator, employer, prior provider, or record-entry method were not analyzed in this summary and are candidates for follow-up analysis. Historical audiograms are an uncontrolled baseline, so this study quantifies real-world agreement between measurement processes, not absolute accuracy.
Conclusion
Across 9,496 untrimmed matched comparisons, Soundtrace PTA measurements agree with historical occupational audiograms with a mean bias of -2.68 dB (95% CI -2.83 to -2.53, p < 0.001) and 91.4% of comparisons within ±10 dB, consistent with published boothless validation studies despite an uncurated, heterogeneous baseline [4, 5]. The mean difference remained relatively stable across the untrimmed and percentile-restricted sensitivity analyses while dispersion contracted, indicating a small, systematic, clinically minor offset and a tight central agreement between the two measurement processes.
Given the scale of the dataset, spanning 9,556 employees, more than 300 Soundtrace audiometers, more than 275 Soundtrace certified facilitators, and more than 250 employers, the results provide a stable estimate of real-world PTA agreement and support the combination of calibration management, boothless audiometry with enforced permissible noise levels, and standardized facilitator training.
Key Findings
Sources & References
- 1.[1] Mahomed F, Swanepoel DW, Eikelboom RH, Soer M. Validity of automated threshold audiometry: a systematic review and meta-analysis. Ear and Hearing. 2013;34(6):745-752.
- 2.[2] Schmuziger N, Probst R, Smurzynski J. Test-retest reliability of pure-tone thresholds from 0.5 to 16 kHz using Sennheiser HDA 200 and Etymotic Research ER-2 earphones. Ear and Hearing. 2004;25(2):109-119.
- 3.[3] Dobie RA. Reliability and validity of industrial audiometry: implications for hearing conservation program design. The Laryngoscope. 1983;93(7):906-927.
- 4.[4] Swanepoel DW, Mngemane S, Molemong S, Mkwanazi H, Tutshini S. Hearing assessment: reliability, accuracy, and efficiency of automated audiometry. Telemedicine and e-Health. 2010;16(5):557-563.
- 5.[5] Brennan-Jones CG, Eikelboom RH, Swanepoel DW, Friedland PL, Atlas MD. Clinical validation of automated audiometry with continuous noise-monitoring in a clinically heterogeneous population outside a sound-treated environment. International Journal of Audiology. 2016;55(9):507-513.
- 6.ANSI S3.6-2018: Specification for Audiometers
- 7.ANSI S3.1-1999 (R2013): Maximum Permissible Ambient Noise Levels for Audiometric Test Rooms
- 8.OSHA 29 CFR 1910.95: Occupational Noise Exposure Standard
- 9.ISO 8253-1: Acoustics, Audiometric test methods, Part 1: Pure-tone air and bone conduction audiometry