August 8, 2026 · The Key Bot
Measuring Safety Performance in Your Own Traffic Control Company
National work zone statistics describe a country, not your operation. Which metrics actually tell you whether your company is getting safer, how to collect them without a research budget, and how to present them to an agency or an underwriter.

In-depth guide · sources linked inline
Every traffic control company says safety is a priority. Very few can answer a simple follow-up: how would you know if it were getting worse?
That question is harder than it looks, because the obvious answer — we haven't had an injury — is a measurement of luck at the sample sizes most companies operate at. A twenty-crew company can go a year without a recordable injury while running a genuinely deteriorating operation, and it can have a serious injury in a year where everything was done well.
National statistics do not solve this. They describe a population, and the ones worth knowing are covered in work zone safety statistics explained. This article is about the other half of the problem: what you can measure about your own company, and what to do with it.
Why the obvious metrics are not enough
The metrics companies report externally — recordable incident rates, lost-time rates, EMR — share a structural weakness for internal management. They are lagging, they are rare-event driven, and at a single company's scale they are statistically noisy.
Consider a company with fifty field staff. In a typical year it might have a handful of recordable incidents. A move from three to one is not evidence of improvement, and a move from one to three is not evidence of decline. Both are within the range of ordinary variation. Managing to those numbers means reacting to noise, which produces the familiar cycle of a safety push after a bad quarter and quiet drift after a good one.
That does not make them useless. They are required, they are what other people judge you by, and over enough years the trend is real. But they cannot tell you this month whether the operation is drifting.
For that you need events that happen often enough to form a signal — which means counting the things that precede harm rather than harm itself.
The five leading indicators worth collecting
None of these require a research budget. All of them require a capture mechanism a crew will actually use.
1. Intrusions and near-misses
Every event where a vehicle entered the work space, or came close enough that someone had to react. Date, time, location, road type, setup type, conditions, and a sentence about what happened.
This is the single most valuable dataset a traffic control company can build about itself, and almost none exist, for a predictable reason: the reporting burden falls on a crew at the end of a long day, and nothing happened, so why write it up.
Two things make it work. First, the report has to take under a minute, on a phone, at the site. Second, and more important, reporting has to be visibly safe. If the first detailed near-miss report produces an investigation into whether the crew did something wrong, you will not get a second one. The correct response to a near-miss report is thanks, followed by a question about the site.
Expect the count to rise in year one. That is the program working. Say so out loud, repeatedly, or people will read the rise as evidence they are being blamed for something.
2. Device strikes
Which devices were hit, where, under what conditions. Cheap to collect because someone already has to replace the device, and revealing over a season: strikes cluster at particular setup patterns, particular roads, and particular times, which is a direct map of where drivers are being surprised.
It also feeds procurement, since a device destroyed is a device to replace — see charging for damaged and lost traffic control devices.
3. Setup and takedown counts
If exposure concentrates when devices are being placed and removed — and both practitioner experience and the structure of the standards point that way — then the number of setups is a direct exposure measure.
This one is quietly important because it decouples exposure from revenue. Two companies with identical revenue can differ substantially in setups performed, and the one doing more short jobs is carrying more exposure for the same money. That is a pricing insight as much as a safety one, and it argues for taking short-duration and mobile operations seriously as a distinct category rather than as small versions of normal jobs.
4. Hours worked, by job type
The denominator for everything else. Without it you have counts, and counts cannot distinguish "we got worse" from "we did more work."
Most companies have this already in payroll but not attributed to job type, which is where it becomes useful. Hours on high-speed facilities are not equivalent to hours on residential streets.
5. Inspection deficiencies, by category
Both your own self-inspections and agency findings, categorized rather than merely counted — see work zone inspections and agency audits.
The categorization is the whole point. Fifteen deficiencies spread across fifteen categories is ordinary variation. Fifteen deficiencies of which nine are advance-warning spacing is a training gap with an address. Individual findings feel like carelessness; the category distribution shows you it is systematic.
The lagging indicators you will be asked for
You still need these, because other people use them as gates.
TRIR — Total Recordable Incident Rate — is recordable incidents × 200,000 ÷ hours worked, where the multiplier normalizes to 100 employees at 2,000 hours a year. What counts as recordable is determined by OSHA's recordkeeping regulation at 29 CFR Part 1904, with the agency's practical material collected on its recordkeeping page. Determine recordability against the regulation rather than by instinct — under-recording creates a compliance exposure, and over-recording inflates a number that gates your bidding.
DART — days away, restricted, or transferred — uses the same formula over the subset of incidents involving those outcomes. It is generally a better severity signal than TRIR alone.
EMR — the experience modification rate — comes from the workers' compensation system rather than from you, calculated by a rating bureau such as NCCI or a state equivalent from your claims history against expected experience for your classification. Clients like it precisely because you did not calculate it.
Three things about EMR worth internalizing, because it is frequently a hard gate: it is slow, lagging several years, so today's improvements show up much later; it is claims-driven, meaning claim management practice affects it independently of incident frequency; and it is commonly used as a pass/fail threshold in prequalification regardless of context. See safety records and EMR in traffic control prequalification and DOT prequalification.
For industry context on the lagging side, the Bureau of Labor Statistics Census of Fatal Occupational Injuries publishes the fatal injury tables, which have reported total U.S. workplace fatalities above 5,000 per year across all industries, with transportation incidents persistently among the largest single event categories. The Injuries, Illnesses, and Fatalities program also publishes non-fatal rates by industry, which is where a benchmark for your own TRIR should come from rather than from a competitor's marketing.
Two more things worth counting, once the basics run
Once the five leading indicators are being collected reliably, two additions repay the effort.
Briefings held versus jobs run. A simple ratio, and an uncomfortable one the first time you calculate it. Most companies believe tailgate briefings happen on essentially every job; the record usually shows otherwise, particularly on short jobs, second setups of the day, and anything running late. Because the briefing is where site-specific hazards get raised, gaps in it correlate with the jobs where nobody thought about the site. See tailgate safety meetings that crews actually use.
Time-of-day and light-condition distribution of your work. Night work carries a different risk profile from daytime work, and a company whose night share has grown from a tenth to a third has materially changed its exposure without anyone deciding to. This is worth knowing before an underwriter points it out, and it is worth pricing — see night work traffic control.
Both are derived from data you already have if jobs, crews, and times are recorded. Neither requires anyone to fill in anything new, which is why they are the right second step rather than the first.
Where the regulatory floor sits
None of this replaces the compliance baseline, and it is worth being clear which is which so the two do not get conflated in a proposal.
Employer obligations in road work come from OSHA, whose highway work zones material collects the guidance, with the construction standard at 29 CFR Part 1926, Subpart G, which OSHA describes as covering "Signs, Signals, and Barricades", and flagging provisions at 29 CFR 1926.201.
The traffic-guidance standards are separate and come from the Manual on Uniform Traffic Control Devices, currently the 11th Edition issued in December 2023, as adopted or supplemented by your state — with requirements varying by state, county, and city, and the authority having jurisdiction as the only definitive source for a given site.
On federal-aid projects, agencies operate under 23 CFR Part 630, Subpart J, titled "Work Zone Safety and Mobility", which is where a good deal of the specification language you comply with originates.
Measurement sits above that floor. Compliance tells you whether a setup met a standard; measurement tells you whether your company is getting better at meeting it. Presenting the first as though it were the second is the most common weakness in contractor safety narratives — a list of requirements met is not evidence of a program, because meeting requirements is the minimum condition of operating at all.
Collecting it without a research budget
The entire practical difficulty is capture. Every metric above is easy to define and hard to collect, for the same reason: the person who has the information is standing at a roadside at the end of a shift.
Four design rules that determine whether a program produces data or produces good intentions.
Capture at the point of occurrence. A form completed at the office is completed from memory, late, and incompletely. The gap between site capture and office capture is larger than any analytical technique could compensate for.
Pre-populate everything you already know. Date, time, location, job, crew. If a foreman has to type what the system already knows, the friction is self-inflicted.
Under a minute, or it will not happen. This is not a target, it is a threshold. Above it, capture rate falls off sharply, and the fall is worst on exactly the busy days you most want data from.
One place, not five. Briefings, tickets, inspections, intrusions, and device movements captured in five separate systems means four of them decay. They belong on the job record together, which is also what makes them analyzable together later.
The reason to be strict about this is that the failure is invisible. A program with 30% capture does not announce itself — it produces a clean-looking dataset showing a safe operation, because the events that went unrecorded are disproportionately the ones from the worst days.
Reading the numbers once you have them
A few habits keep the analysis honest.
Rates, not counts, for anything you compare across periods. Normalize by hours or by setups. A quarter with more incidents and much more work may be an improvement.
Look at distributions, not just totals. Where are deficiencies concentrated — which crew, which road type, which time of day, which job type? Totals hide everything actionable.
Treat a rising near-miss count as good news until proven otherwise. In a young program it almost always means reporting improved.
Watch the ratio of self-found to agency-found deficiencies. A high proportion found by your own inspections means the program is working. A high proportion found by inspectors means it is not, regardless of the total.
Do not set targets on lagging indicators for crews. Injury-rate targets attached to individuals or crews reliably suppress reporting rather than injuries. Set targets on leading activities — briefings held, inspections completed, near-misses reported — where more is unambiguously better.
Presenting it externally
This is where the effort pays for itself commercially, and most contractors underuse what they have.
Agencies and underwriters are trying to distinguish companies with a functioning safety program from companies with a safety binder. Almost every bidder asserts the former. Very few can demonstrate it.
What demonstrates it: a trend line with an explanation, a category breakdown with an intervention attached, and evidence that something changed as a result. "Advance-warning spacing was our largest deficiency category last spring; we changed the release format to carry spacing for the road class, retrained two crews, and it fell to third" is worth more than any statement of commitment. It shows a loop — measure, find, act, re-measure — which is exactly the thing being assessed.
Two cautions. Present your own data honestly including the bad periods, because a suspiciously clean record reads as a reporting problem to anyone experienced. And keep national context in proportion: cite the National Work Zone Safety Information Clearinghouse or FHWA's work zone facts and statistics as framing, not as your evidence. The frame is the country; the evidence is you.
Starting from nothing
If none of this exists today, the order that works:
Month one — intrusions and near-misses only. One metric, captured well, with a visible no-blame response. Do not add anything else until reporting is actually happening.
Month two — inspection deficiencies, categorized. You are probably already inspecting; the change is recording the category.
Month three — hours by job type and setup counts. Mostly a matter of attributing data you already collect.
Month four — review. Look at distributions, pick one intervention, make it, and keep measuring.
Resist the temptation to launch all of it at once with a kickoff meeting and a binder. Safety measurement programs fail the same way software implementations do: a broad launch that stumbles becomes a story people tell for years about how the whole idea did not work, while a narrow one that succeeds becomes the thing crews ask to extend. One metric, collected well, with a visible response, is worth more than six collected indifferently.
The other common failure is starting with the metric that is easiest to collect rather than the one that matters. Hours worked is trivially available and tells you almost nothing on its own; near-misses are the hardest to collect and carry nearly all the signal. Start with the hard one while the program has attention and goodwill behind it.
Nothing above is sophisticated. The difficulty is entirely in sustaining capture across crews on busy days, which is a tooling problem more than a discipline problem. That is the specific thing Traffic OS is built for in this industry — briefings, inspections, intrusions, device movements, and GPS-stamped signed tickets captured on the job from the field, on flat-tier pricing by company size rather than per user, because a safety dataset with the seasonal crews missing is not a dataset. If you cannot currently answer how many intrusions you had last quarter, book a walkthrough — that question is usually where this starts.
Frequently asked questions
What safety metrics should a traffic control company track?+
A workable set is: intrusions and near-misses, device strikes, setup and takedown counts, hours worked by job type, inspection deficiencies by category, and the standard recordable-injury rates your insurer and prospective clients will ask for anyway. The first five are leading indicators you can act on; the recordable rates are lagging indicators other people use to judge you.
What is the difference between a leading and a lagging safety indicator?+
A lagging indicator counts harm that already happened — injuries, claims, incident rates. A leading indicator counts conditions and behaviours that precede harm — near-misses, deficiencies found, briefings held, exposures created. Lagging indicators are what get reported externally; leading indicators are what actually let you change an outcome, because by the time a lagging indicator moves, the event has occurred.
How is TRIR calculated?+
Total Recordable Incident Rate is the number of OSHA-recordable incidents multiplied by 200,000, divided by total hours worked. The 200,000 represents 100 employees working 2,000 hours a year, which normalizes companies of different sizes. What counts as recordable is defined by OSHA's recordkeeping regulation, not by internal judgment, and getting that determination wrong in either direction causes problems later.
What is EMR and why does everyone ask for it?+
The experience modification rate is a workers' compensation underwriting factor comparing your claims history to the expected experience for your classification and payroll. Clients and agencies use it as a convenient third-party safety proxy because it is externally calculated and hard to manipulate. It is a lagging indicator, it is slow to move, and it says less about current practice than people assume — but it is frequently a hard gate on bidding.
How many near-misses should we expect to record?+
Far more than you currently do. A company recording almost no near-misses is not safer than one recording many — it has a reporting problem. Rising near-miss counts in the first year of a new program are the program working, and it is worth telling the crews that explicitly, or the numbers will go back down.
Can we use our own data with agencies and insurers?+
Yes, and it is usually more persuasive than anything else you can offer. A contractor who can show intrusion trends, inspection findings by category, and what changed as a result is demonstrating a functioning program rather than asserting one. That is exactly the distinction prequalification reviewers and underwriters are trying to make.