Acoustic Gunshot Detection vs. AI Visual Gun Detection: The 2026 Threat Intelligence Briefing on the Field-Confirmation Gap and the Post-Shot/Pre-Shot Divide
A primary-source analysis of why acoustic gunshot detection and AI visual gun detection answer different questions, what the Chicago and New York audits revealed about field confirmation, and how the post-shot/pre-shot divide should drive the buying decision.
Acoustic gunshot detection and AI visual gun detection are not two products competing for the same job. They are two different jobs that buyers keep confusing for one, and the confusion is expensive. One listens for the sound a firearm makes after it is fired and races to put officers and medics near the gunfire faster than a 911 caller could. The other watches for the firearm itself, in the seconds before the trigger is pulled, so that a building can be locked down while the weapon is still in a hallway rather than after the first round has already left it. The split between post-shot and pre-shot is the single most important distinction in this category, and almost no procurement document draws it cleanly.
This briefing is for school safety directors, hospital security leaders, municipal risk officers, and the procurement teams who have to decide what a gunshot-detection line item is actually buying. It works through the physics of acoustic detection, the field-confirmation gap documented by the City of Chicago and the City of New York in their own audits, the independent evidence that acoustic systems do deliver real value on the metrics they were designed for, and the architectural reason a visual detection layer answers a question the acoustic layer cannot. The goal is not to declare a winner. It is to make the post-shot and pre-shot divide legible so that a buyer chooses the modality that matches the threat they are actually trying to interrupt.
Three numbers from government audits define the gap between what acoustic gunshot detection promises and what it delivers in the field.
Each of those numbers comes from a primary source, and together they frame the analysis. The 9.1 percent figure is from the Chicago Office of Inspector General, which found that of 41,830 ShotSpotter alerts with a recorded disposition, only 4,556 led to evidence of a gun-related criminal offense. The 87 percent figure is from the New York City Comptroller's 2024 audit, which reported that ShotSpotter alerts were confirmed as actual shootings only 13 percent of the time. The zero is physics: a microphone cannot hear a shot that has not happened. Holding those three facts together is what separates a clear-eyed evaluation from a marketing comparison, and it is the work this briefing does.
The distinction matters because the threat-intelligence frame for any detection system is the same one IntelliSee applies across the category: evaluate the system against the conditions it will actually face and the question it is actually being asked to answer, not the demonstration it was sold under. Our companion briefing on AI gun detection failure modes makes the case that a detection layer has to be judged on its real architecture gaps. Acoustic detection is the same problem heard rather than seen, and its architecture gap is a timing gap.
How acoustic gunshot detection actually decides a shot was fired
Acoustic gunshot detection works by listening for two distinct sound events and using their arrival times at multiple sensors to locate the source. Understanding that mechanism is what makes the modality's strengths and its confirmation gap both legible.
When a firearm discharges, it produces a muzzle blast: an explosive pressure wave created by the rapidly expanding propellant gases escaping the barrel, radiating outward at the speed of sound. If the bullet travels faster than sound, it also produces a ballistic shockwave, a bow-shaped pressure front trailing the projectile with a characteristic "N-wave" signature whose shape encodes information about caliber and trajectory. A network of acoustic sensors, typically mounted on rooftops and utility poles across a coverage area, records the time of arrival of these wavefronts. Because the sensors sit at known locations, the system computes the time difference of arrival between sensors and uses multilateration, the same geometric principle behind GPS, to triangulate the origin of the sound. These mechanics are well documented in the patent and government-research literature, including the National Institute of Justice's foundational report on gunshot detection systems.
That architecture has a genuine virtue. It is fast at the thing it does. A multi-sensor acoustic array can register a candidate gunshot and publish a located alert within seconds, often before any human has dialed 911, and it does not depend on a witness being present, willing, or accurate about the address. Those are real advantages, and the independent evidence that they hold up in the field is covered later in this briefing. But the architecture also fixes the modality on one side of the event. Every signal it processes is generated by a shot that has already been fired. The muzzle blast is the trigger. There is no acoustic signature for a gun being carried, raised, or aimed, because none of those actions makes a sound. This is not a limitation that better algorithms can remove. It is the defining property of the modality.
The field-confirmation gap: what the Chicago and New York audits actually found
The most important number in acoustic gunshot detection is not its detection accuracy. It is the share of its alerts that turn out to correspond to a confirmed shooting, and two of the largest deployments in the country have had that number measured by their own governments.
In Chicago, the Office of Inspector General analyzed every ShotSpotter alert between January 2020 and May 2021. Of 50,176 alerts confirmed as probable gunshots and dispatched, 41,830 had a recorded disposition. Of those, 4,556 produced evidence of a gun-related criminal offense: 9.1 percent. The OIG's conclusion was carefully worded and worth quoting in spirit: police responses to the alerts rarely produced evidence of a gun crime and seldom led to investigatory stops with demonstrable value. A parallel study by the MacArthur Justice Center at Northwestern, reviewing roughly 21 months of the same city's data, found that 89 percent of alerts turned up no gun-related crime and 86 percent no report of any crime at all, which it characterized as more than 40,000 dead-end deployments.
New York's audit told a structurally identical story with a different agency holding the pen. The NYC Comptroller's June 2024 review found that ShotSpotter alerts were confirmed as actual confirmed shootings only 13 percent of the time. The other 87 percent sent officers toward loud noises that did not turn out to be confirmed gunfire, amounting to 7,262 unconfirmed incidents across the sampled months. The audit also quantified the operational drag: officers spent an average of 20 minutes investigating alerts deemed unfounded and 32 minutes on alerts that simply went unconfirmed, time pulled from a finite patrol capacity.
It is important to be precise about what these audits do and do not say, because the vendor and its critics measure different things, and both are partly right. SoundThinking, the company behind ShotSpotter, reports a technical detection accuracy above 90 percent, a figure an independent Edgeworth Economics audit supported for detecting, classifying, and publishing gunfire incidents. That is a measure of whether the system correctly identifies the acoustic event it heard. The government audits measure something else: whether responding officers found confirmable evidence of a crime at the scene. The gap between those two numbers is the field-confirmation gap, and it is the number that governs whether a buyer is paying for outcomes or for alerts. A system can be technically accurate at hearing gunshots and still send officers to a large majority of scenes where nothing prosecutable is found, because gunfire is not always a crime, evidence is not always recoverable, and not every loud impulsive sound is a gun.
Two Measurements, Two Very Different Numbers
| What is being measured | The number | Source and meaning |
|---|---|---|
| Technical detection accuracy | 90%+ | Vendor and independent audit (Edgeworth). Did the system correctly detect, classify, and publish the acoustic gunfire event it heard? |
| Chicago field confirmation | 9.1% | Chicago OIG, 41,830 dispositions. Share of police responses that found evidence of a gun-related crime. |
| Chicago dead-end deployments | 89% | MacArthur Justice Center, ~21 months. Share of alerts turning up no gun-related crime. |
| New York confirmation rate | 13% | NYC Comptroller, 2024. Share of alerts confirmed as actual shootings. |
| NYPD officer time per unconfirmed alert | 20–32 min | NYC Comptroller, 2024. Average investigation time per unfounded or unconfirmed alert. |
The point is not that acoustic detection is fraudulent. It is that the headline accuracy number and the field-confirmation number answer different questions, and a buyer who reads the first as if it were the second will be surprised by what arrives after go-live. This is the same evaluation discipline our procurement and proof-of-concept methodology for AI gun detection applies to the visual side: insist on the metric that maps to your outcome, not the metric that demos well.
What acoustic detection does deliver, measured independently
A briefing that only catalogued the confirmation gap would be making the same error it warns against: judging a modality on one number. Acoustic detection has a legitimate job, and the strongest independent evaluation to date found that it does that job.
In August 2025, researchers at Cleveland State University published a 185-page evaluation of the Cleveland Division of Police's ShotSpotter deployment, commissioned by the city and conducted across all five police districts from the summer of 2023 through the summer of 2025. The study analyzed roughly 87,000 alerts, observed how officers used the technology in practice, and surveyed both police and residents. Its findings were mixed in a way that is genuinely useful. The system accurately detected and located gunfire, alerted police within seconds, and was more reliable at directing officers to the actual location of gunfire than 911 callers were, including for gunshots that residents never reported at all. That faster, more accurate routing helped officers reach victims and evidence sooner, which the researchers connected to lives saved that might otherwise have been lost.
The same study was equally clear about the limits. ShotSpotter did not, on its own, reduce crime, and the volume of alerts added roughly 20 additional priority-one calls per day, straining police resources in exactly the way the New York audit described. The honest reading of the Cleveland evidence is that acoustic detection is a response-acceleration tool, not a crime-prevention tool, and that its value is concentrated entirely after a shot is fired. It gets help to a shooting scene faster. It does nothing to stop the shooting from starting. That is not a criticism of the technology. It is a description of which half of the timeline it operates on.
Where each detection modality lives on the attack timeline
Acoustic detection begins working at the instant of the first shot. Visual detection begins working when the weapon enters the frame, which is earlier. The two layers cover different seconds.
Post-shot detection and pre-shot detection are not substitutes
The reason buyers conflate these modalities is that both have the word "gun" in the category name. But an acoustic array and a visual gun-detection layer occupy different seconds of the same event. Acoustic detection begins at the muzzle blast and accelerates the response to a shooting that is already underway, which is real and valuable for getting medics to a bleeding victim inside the survival window. Visual detection begins when the weapon becomes visible to a camera, which can be before the first shot, and its job is to compress the interval between a weapon appearing and a building responding. A site that wants faster medical and law-enforcement response to gunfire and a site that wants to interrupt an attack before it starts are asking for different things. The mistake is treating one modality as a cheaper version of the other.
Why the post-shot window is a medical clock, not a security clock
The strongest case for acoustic detection is medical, and understanding it clarifies exactly what the modality can and cannot buy. Once a person has been shot, survival is governed by a clock measured in minutes, and acoustic detection competes against that clock.
Trauma medicine is blunt about the timeline. A victim with a femoral-artery injury can bleed out in under three minutes, and severe hemorrhage can become unsurvivable inside roughly two minutes without intervention. Tactical-medicine practice describes the first five minutes after a penetrating injury as the decisive window, and the data on tourniquets is stark: troops who had one applied before going into shock survived at a rate near 96 percent, against roughly 4 percent for those treated after shock set in. Against that biology, the average urban ambulance response of about eight minutes is already too slow for an arterial bleed, which is why getting any trained responder to the victim faster is a genuine intervention. An acoustic system that puts officers near a confirmed shooting a minute or two faster than a 911 call is operating directly on this survival clock, and that is the legitimate core of its value.
But notice what that framing concedes. The medical clock only starts after someone has been shot. Acoustic detection optimizes the response to an injury that has already occurred. It is, in the most literal sense, a tool for the aftermath. The security question that comes before the medical one, can the attack be interrupted before the first round is fired, is not a question acoustic detection is built to answer, because the modality has no signal until the shot exists. This is the same response-versus-prevention tension that runs through IntelliSee's ROI framework on detection-to-response latency economics: compressing the seconds after an event is worth real money and real lives, and compressing them to zero by detecting before the event is worth more.
The visual layer answers the question acoustic cannot
Visual gun detection occupies the seconds that acoustic detection cannot reach, because it keys on the weapon rather than the gunshot. That single architectural difference is the whole of the post-shot/pre-shot divide.
An AI visual detection layer runs computer-vision models against existing camera feeds and searches for the object: a firearm being carried, drawn, or raised. When a weapon enters a camera's field of view, the model can flag it with a confidence score and route an alert, as in the entrance-camera detection shown earlier in this briefing, where a weapon was identified at 0.91 confidence before any shot was fired. The mechanics of how those models are trained and where they fail are covered in depth in our technical reference on how AI gun detection works. The relevant point for a modality comparison is timing: the weapon is visible before it is fired, so a system watching for the weapon has a window that a system listening for the shot structurally does not.
That window is where interruption becomes possible. A weapon detected at a building entrance can trigger a coordinated response while the firearm is still in a vestibule: doors lock, a mass-notification message goes out, and a verified alert with a camera location reaches a public-safety answering point, all before the attacker reaches an occupied space. IntelliSee's architecture for this is detailed in our briefing on detection-to-lockdown architecture, which connects the detection event to access control, notification, and dispatch. None of that is available to a modality whose earliest possible signal is the first gunshot, because by then the interruption window has already closed.
There is also a verification dimension that matters for the false-confirmation problem the audits exposed. A visual detection carries its own ground truth: the alert is attached to an image of the object that triggered it, which a human can confirm or dismiss within seconds. An acoustic alert, by contrast, is an inference from sound with no visual to check against, which is part of why officers spent 20 to 32 minutes per unconfirmed New York alert establishing what, if anything, had happened. The presence of a confirming image is exactly what our analysis of swatting and the ground-truth detection layer identifies as the difference between an alert that consumes response capacity and one that directs it. IntelliSee pairs detection with human-in-the-loop confirmation so that a flagged weapon is verified before it escalates, keeping the false-positive cost bounded while preserving the pre-shot timing advantage.
Detecting the weapon, not the person, is what keeps a visual layer deployable
A visual gun-detection layer raises an obvious privacy question, and the architecture answer is what determines whether it clears review. IntelliSee detects the object, a firearm in the frame, rather than identifying the individual carrying it. It performs no facial recognition, stores no video, and collects no protected health information. Detection runs on an on-premises appliance, so feeds do not leave the local network, and the inference is about what is present in the scene, not who is present. That is the architectural reason a weapon-detection layer can deploy across schools, hospitals, and public buildings without inheriting the identity-surveillance objections that a facial-recognition system would face. It is also a meaningful contrast with the acoustic modality, whose civil-liberties scrutiny has centered on where its sensors are placed and what else those microphones might capture.
Which modality fits which site
The right modality follows from the threat a site is trying to interrupt and the response capacity it has to absorb alerts. Four deployment profiles cover most real decisions.
Wide outdoor municipal areas
For a police department trying to locate gunfire across square miles of streets where cameras cannot cover everything, acoustic detection has a real role: it surfaces unreported gunfire and routes responders faster than 911. The buyer should go in clear-eyed about the field-confirmation gap and the officer-time cost, and should treat it as a response tool, not a prevention tool.
Defined facilities with cameras
For a school, hospital, or corporate campus with an existing camera footprint and a defined perimeter, visual detection is the stronger fit. The threat to interrupt is an armed person entering a building, the cameras already exist, and the pre-shot window is where lockdown and notification actually save lives. Acoustic coverage indoors is a poor match for the geometry.
High-stakes interruption mandates
Where the explicit goal is to stop an attack before the first shot, only a pre-shot modality qualifies. A post-shot system cannot meet an interruption mandate, by definition. Sites with active-assailant response obligations under standards like NFPA 3000 need detection that fires before gunfire, which is the visual layer's window.
Constrained response capacity
Where the response team is small and cannot absorb a high false-alert load, the field-confirmation rate becomes decisive. A modality that sends responders to a large majority of scenes with nothing to find will exhaust a lean team. Image-confirmable alerts that a human can clear in seconds protect scarce response capacity better than sound-only inferences.
For most defined facilities, the practical answer is not acoustic versus visual as a binary. It is recognizing that a camera-based site already owns the infrastructure for pre-shot detection and should not buy a post-shot tool expecting it to do a pre-shot job. The IntelliSee capability set for the visual layer, including VMS integration, on-premises processing, and human-confirmation routing, is detailed on the weapon detection solution page, and the broader platform context sits on the how it works overview.
Frequently asked questions about acoustic vs. visual gun detection
What is the difference between acoustic gunshot detection and AI visual gun detection?
Acoustic gunshot detection listens for the sound of a fired weapon and triangulates the location of the gunfire from multiple sensors, so it operates entirely after the first shot. AI visual gun detection runs computer-vision models on camera feeds to identify the firearm itself, which can happen before the trigger is pulled because a weapon is visible before it is fired. The core distinction is post-shot versus pre-shot: one accelerates the response to a shooting already underway, while the other aims to surface the weapon in time to interrupt the attack. They occupy different seconds of the same event and are not interchangeable.
How accurate is ShotSpotter, and why do city audits report low numbers?
The discrepancy comes from measuring two different things. The vendor and an independent Edgeworth Economics audit report technical detection accuracy above 90 percent, meaning the system correctly detects and classifies the acoustic gunfire event it heard. Government audits measure field confirmation instead: the Chicago Inspector General found that 9.1 percent of police responses to alerts produced evidence of a gun-related crime, and the New York City Comptroller found alerts were confirmed as actual shootings 13 percent of the time. Both measurements can be accurate because gunfire is not always a recoverable crime, and not every impulsive sound is a gun. Buyers should ask which number maps to the outcome they are paying for.
Does acoustic gunshot detection prevent shootings?
No, and the strongest independent evidence is explicit on this point. The Cleveland State University evaluation published in 2025, which analyzed roughly 87,000 alerts over two years, found that ShotSpotter accurately located gunfire and helped officers reach victims faster than 911 callers, but did not on its own reduce crime. Because the modality has no signal until a shot is fired, it cannot interrupt an attack before it begins. Its documented value is response acceleration after gunfire, which is real and can save lives in the medical survival window, but it is a post-event tool rather than a prevention tool.
Can a microphone-based system detect a gun before it is fired?
No. Acoustic detection keys on the muzzle blast and, for supersonic rounds, the ballistic shockwave, both of which are produced by a shot that has already been fired. Carrying, drawing, or aiming a firearm produces no sound, so there is no acoustic signature to detect in the pre-shot window. This is a property of the physics, not a limitation that better algorithms can overcome. Detecting a weapon before it is fired requires a modality that keys on the visible object, which is what camera-based visual detection does.
Is acoustic or visual gun detection better for a school or hospital?
For a defined facility with an existing camera footprint and a perimeter to protect, visual detection is generally the stronger fit. The threat to interrupt is an armed person entering a building, the cameras already exist, and the pre-shot window is where lockdown and mass notification actually change the outcome. Acoustic systems are designed for wide outdoor municipal coverage and are a poor geometric match for indoor facility protection. A school or hospital trying to interrupt an attack before the first shot needs a pre-shot modality, which acoustic detection cannot provide.
Does IntelliSee use facial recognition or store video for gun detection?
No. IntelliSee detects the weapon as an object in the camera frame rather than identifying the person carrying it. It performs no facial recognition, stores no video, and collects no protected health information. Detection runs on an on-premises appliance so feeds do not leave the local network, and the inference is about what is present in the scene, not who is present. This architecture is what allows a visual weapon-detection layer to deploy across schools, hospitals, and public buildings without the identity-surveillance objections a facial-recognition system would raise.
How does IntelliSee avoid the false-alert problem the city audits described?
By pairing detection with human-in-the-loop confirmation and by attaching a visual to every alert. Because a visual detection is tied to an image of the object that triggered it, a person can confirm or dismiss it within seconds, rather than dispatching responders to a sound-only inference that takes 20 to 32 minutes on scene to establish what happened. Detection thresholds and zones are tunable, and a flagged weapon is verified before it escalates to lockdown or dispatch. The goal is to preserve the pre-shot timing advantage while keeping the false-positive cost bounded enough that a lean response team can sustain it.
Continue the research
This briefing covers the physics of acoustic detection, the field-confirmation gap documented by the Chicago and New York audits, the independent evidence of what acoustic detection does deliver, and the architectural reason a visual layer answers a question acoustic cannot. For deeper reading on the connected questions:
- How AI Gun Detection Works: A Technical Reference — the computer-vision architecture, model training, and accuracy trade-offs behind the visual modality.
- Detection-to-Lockdown Architecture — how a pre-shot detection event connects to access control, mass notification, and PSAP dispatch to close the action gap.
- AI Gun Detection Failure Modes — the threat-intelligence analysis of what visual computer vision misses and why, so the modality is evaluated honestly.
- Weapon detection solution page — the IntelliSee capability set, including VMS integration, on-premises processing, and human-confirmation routing.
If you are weighing a gunshot-detection line item and want to map the post-shot and pre-shot question to your own site before you commit, a structured risk assessment is the place to start.
More intelligence like this
New IntelliSee research drops monthly at most. Subscribe and get the next sector playbook, technology briefing, or threat intelligence report in your inbox the day it ships.
Request a Risk Assessment
Talk to an IntelliSee security specialist. No sales pitch — a structured conversation about your environment, your threat profile, and whether computer vision is the right fit.
Request a Risk Assessment