Camera Requirements for AI Video Analytics: A 2026 Technology Briefing on the DORI Standard, Johnson Criteria, Pixel Density Math, and the Imaging Variables That Determine Detection Performance
Home / Intelligence / Camera Requirements for AI Video Analytics:...
Technology Briefings

Camera Requirements for AI Video Analytics: A 2026 Technology Briefing on the DORI Standard, Johnson Criteria, Pixel Density Math, and the Imaging Variables That Determine Detection Performance

A 2026 technology briefing on the DORI standard, Johnson criteria, pixel density math, and the imaging pipeline variables that determine whether AI detection performs on the cameras you already own.

Published June 2026
Read Time 16 min read
Stream Technology Briefings
250 px/m
Pixel density IEC 62676-4 anchors to identification tasks, ten times the 25 px/m detection floor (IEC 62676-4:2014)
~20%
Benchmark mean average precision for small objects under 32 pixels, less than half large-object accuracy (peer-reviewed detection surveys)
1B+
Surveillance cameras in the global installed base, most deployed before AI readiness was a design criterion (Omdia-aligned industry estimates)

Three numbers explain why AI video analytics succeed on some camera estates and quietly underperform on others: the pixel density the international standard demands, the accuracy penalty detection models pay on small targets, and the size of the installed base that was never designed with either in mind.

250 px/mPixel density IEC 62676-4 anchors to identification tasks, ten times the 25 px/m detection floor (IEC 62676-4:2014)
~20%Benchmark mean average precision for small objects under 32 pixels, less than half large-object accuracy (peer-reviewed detection surveys)
1B+Surveillance cameras in the global installed base, most deployed before AI readiness was a design criterion (Omdia-aligned industry estimates)

Camera requirements for AI video analytics come down to one question that every procurement eventually collides with: will the model perform on the cameras we already own? Security directors evaluating detection platforms tend to scrutinize the AI vendor's accuracy claims and ignore the variable that more often decides field performance, which is the imaging chain feeding the model. A detection model never sees a hallway, a parking lot, or a drawn firearm. It sees a grid of pixels, and the number of those pixels that land on the target is governed by physics and geometry that were standardized decades before deep learning existed.

This briefing is a technical reference for security directors, integrators, and procurement teams who need to assess camera requirements for AI video analytics before a deployment. It covers the 1958 Johnson criteria that still underpin surveillance performance math, the IEC 62676-4 DORI framework that translates operational tasks into pixel densities, the ways machine detection rewrites those human-operator assumptions, and the pipeline variables beyond resolution, compression, frame rate, and lighting among them, that quietly break detection in the field. It closes with an audit framework for scoring an existing camera estate, because the practical question for most buyers is not which cameras to buy. It is which of the cameras they already own are ready.

The Johnson criteria: the 1958 military research that still governs surveillance math

Nearly every camera performance specification in modern physical security descends from experiments run by a U.S. Army scientist in 1957 and 1958. John Johnson, working at what became the Army's Night Vision and Electronic Sensors Directorate, used image intensifier equipment and volunteer observers to measure how much visual information a human needs to perform four tasks against a target: detect that something is present, determine its orientation, recognize what class of object it is, and identify the specific object. He presented the findings at the first Night Vision Image Intensifier Symposium in October 1958 in a paper titled Analysis of Image Forming Systems, and the thresholds it contained became known as the Johnson criteria. The history and evolution of the criteria is documented in U.S. Department of Energy technical literature.

Johnson expressed his thresholds in line pairs, the distance spanned by one dark and one light line at the limit of an observer's acuity, resolved across the target's critical dimension. Detection required roughly 1 line pair. Recognition, telling a person from a car, required about 4 line pairs, plus or minus 0.8. Identification, telling one specific person or vehicle from another, required about 6.4 line pairs, plus or minus 1.5. Two properties of those numbers still matter to every modern buyer. First, they are probabilistic: each threshold represents a 50 percent chance that an observer completes the task, not a guarantee. Doubling the threshold raises task probability toward certainty, which is why field designs build margin above the minimums. Second, the spread between tasks is large. Identification demands roughly six times the resolved detail of detection, and that ratio survived the translation from image intensifier line pairs to digital pixels almost intact.

The Johnson criteria were built for human eyes looking at analog displays, and they assume a trained, attentive observer. Both assumptions break in modern security operations, the first because the consumer of the image is increasingly a neural network, the second because sustained human attention across hundreds of feeds was never realistic. But the core insight, that performance is a function of resolved detail on the target rather than the camera's headline resolution, transferred directly into the digital era's governing standard.

The DORI standard: how IEC 62676-4 anchors operational tasks to pixel density

IEC 62676-4 is the international standard that converts the Johnson framework into pixel arithmetic for digital video surveillance. Published in 2014 as the application guideline volume of the video surveillance standard family, it defines four operational tiers, Detect, Observe, Recognize, and Identify, collectively known as DORI, and anchors each to a minimum pixel density measured on the target plane: 25 pixels per meter for detection, 62.5 px/m for observation, 125 px/m for recognition, and 250 px/m for identification. In imperial terms those are roughly 7.6, 19, 38, and 76 pixels per foot. The densities are derived from how many pixels must land across a human face or body for an operator to complete each task, and the standard's 2025 revision extends the taxonomy further and raises the top-tier guidance to 500 px/m for the most demanding scrutiny tasks.

The arithmetic that matters to a buyer is simple and unforgiving. Pixel density at a point in the scene equals the camera's horizontal resolution divided by the width of the scene the lens covers at that distance. A 1080p camera delivers 1,920 horizontal pixels. Give it an 80 degree lens and point it down a corridor: at 15 meters from the lens, that field of view spans about 25 meters, so the camera is delivering roughly 76 px/m at that depth. That passes the observation tier comfortably and fails recognition. Move the target to 7 meters and the same camera delivers about 163 px/m, clearing recognition with margin. Nothing about the camera changed. The task envelope is a property of geometry, not hardware, and every camera in an estate has a different envelope depending on its lens, mounting height, and the depth of the scene it watches.

This is why megapixel marketing misleads procurement teams. A 4K camera covering a wide parking lot can deliver less density on a target at the far fence line than a 1080p camera with a narrow lens watching a single door. The standard forces the right question, which is not how many pixels the sensor has but how many of them land on the thing you need the system to act on, at the distance where you need the action.

The DORI Ladder

Task-anchored pixel density under IEC 62676-4

Each operational tier demands a minimum density on the target plane. The tier a camera reaches is set by lens, distance, and mounting, not by the megapixel count on the spec sheet.

25 PX/M · 7.6 PX/FT
DETECT

Establish with reasonable certainty that a person or object is present in the scene. The floor for situational awareness and wide-area coverage.

62.5 PX/M · 19 PX/FT
OBSERVE

Characterize what the target is doing: direction of travel, clothing, carried objects, general behavior over time.

125 PX/M · 38 PX/FT
RECOGNIZE

Determine with a high degree of certainty whether an individual has been seen before. The trained-operator tier.

250 PX/M · 76 PX/FT
IDENTIFY

Establish identity beyond reasonable doubt. The evidentiary tier. The 2025 revision raises top-end scrutiny guidance to 500 px/m.

What AI actually needs: why machine detection rewrites the DORI assumptions

DORI was calibrated for human operators, and machine detection does not consume images the way a human operator does. That cuts in both directions, and understanding where it cuts which way is the difference between a camera estate that is genuinely AI-ready and one that merely looks adequate on a coverage map.

The direction that favors the machine: a detection model does not fatigue, does not look away, and evaluates every frame of every feed with the same consistency at hour nine as at minute one. For tasks at the detection and observation tiers, watching a fence line, flagging a person in a restricted zone after hours, registering a vehicle where no vehicle should be, a well-trained model operating on modest pixel densities outperforms any realistic human monitoring arrangement, simply because the human arrangement was never actually watching most feeds. The deeper mechanics of this architecture are covered in the IntelliSee briefing on how AI gun detection works.

The direction that favors caution: small targets are the documented weak point of modern object detection. The research community's standard benchmark protocol scores objects occupying less than 32 by 32 pixels as small, and the performance penalty at that scale is steep and persistent. Peer-reviewed surveys of small-object detection report mean average precision around 20 percent for small objects on the COCO benchmark for mainstream detector families, less than half the accuracy those same models achieve on large objects. The cause is architectural: convolutional networks downsample images as they process them, and a target that begins as a few hundred pixels can shrink below the threshold of representation before the detection layers ever evaluate it. Comprehensive academic reviews describe small-object detection as one of the field's hardest open problems precisely because aggregate accuracy numbers hide it.

Translate that into security terms and the implication is direct. A handgun is a small object at almost any realistic surveillance distance. A firearm that subtends 40 pixels in a frame is a fundamentally easier problem than one that subtends 12, and no amount of model tuning fully erases that gap. Pixel density on target is therefore not a compliance nicety for AI deployments. It is a first-order input to detection probability, which is why a credible vendor conversation starts with your camera geometry rather than with accuracy claims measured on someone else's cameras. It is also why confidence scores matter operationally: a platform that exposes per-detection confidence lets a security team tune alerting thresholds to each camera's actual pixel-density envelope rather than pretending every feed is equal. For the adversarial end of this problem, see the companion analysis of how computer vision models handle occlusion, low light, and adversarial conditions.

Privacy by Design

AI threat detection does not need the identification tier, and that is the point

The top rungs of the DORI ladder exist to answer the question of who someone is. Object-level threat detection answers a different question: what is present and what is happening. Platforms like IntelliSee perform object, posture, and motion-pattern detection, a drawn firearm, a person down, movement into a restricted zone, and do not perform facial recognition, store video, or collect personally identifying biometrics. Architecturally, that means the system operates at the detection and observation tiers where existing camera estates are strongest, rather than demanding the 250 px/m identification densities that drive rip-and-replace projects. The privacy posture and the retrofit economics are the same design decision viewed from two angles.

Real IntelliSee AI gun detection output showing a drawn firearm identified in a 1280x720 camera frame with bounding box and confidence score overlay
LIVE CAM-02 · 1280×720 FEED
Actual IntelliSee detection output. A drawn firearm identified in a 1280×720 camera frame, with the bounding box and confidence score rendered by the platform at the moment of detection. This is the pixel-density argument in a single image: the camera is a standard-resolution feed, not a 4K showpiece, and the firearm is resolved with enough pixels on target for the model to localize and score it. No facial recognition is performed, no video is stored, and the alert reaches designated responders within seconds of the frame being analyzed.

Beyond resolution: the imaging pipeline variables that quietly break detection

Pixel density is the dominant variable in AI detection performance, but it is not the only one, and several of the others are invisible on a camera spec sheet because they are configuration choices rather than hardware properties. Four deserve explicit attention in any readiness assessment.

Compression. Nearly every camera estate compresses video with H.264 or H.265 before the stream reaches anything downstream, and compression destroys exactly the fine detail that small-object detection depends on. The most rigorous public study of this interaction, published on arXiv by researchers testing a YOLO-family detector across 50 annotated surveillance videos, encoded footage at five quality levels and found that detection is generally robust to moderate compression, with constant rate factor settings around 37 preserving performance while sharply cutting bitrate. The same study found performance breaks down at aggressive compression precisely where security needs it most: scenes with challenging illumination, and small, low-contrast, or fast-moving targets. The operational lesson is that bandwidth-saving compression presets configured years ago for storage economics may be silently taxing today's detection accuracy, and that the analytics-facing stream deserves its own quality profile where the VMS supports multi-streaming.

Frame rate. Detection models evaluate individual frames, so analytics do not need cinema-smooth video, and most platforms sample feeds at intervals rather than consuming every frame. What frame rate does govern is the number of opportunities the system gets to see a fast-moving or briefly visible target. A firearm drawn and reholstered in two seconds appears in a handful of analyzable frames at low sample rates. Frame rate also interacts with compression: many encoders hold bitrate constant, so raising frame rate without raising bitrate degrades per-frame quality.

Low light and infrared. Sensor noise rises as light falls, and noise is adversarial texture for a detection model. Cameras with usable IR illumination preserve object silhouettes and edges in darkness, which is what object-level detection actually consumes. The practical audit question is not whether a camera technically produces an image at night but whether the image retains enough contrast structure for edges to survive. Scene lighting upgrades are often the cheapest detection-accuracy improvement available to a facility.

Dynamic range and mounting geometry. Entrances are simultaneously the highest-value detection zones and the worst imaging environments, because doorways backlight every subject who walks through them. Wide dynamic range processing recovers detail in those scenes; cameras without it deliver silhouettes. Mounting geometry compounds the problem: steep downward angles foreshorten people and carried objects, shrinking the effective pixel footprint of exactly the targets that matter. A camera mounted high for vandal resistance may be geometrically incapable of presenting a firearm to the model at a usable aspect, regardless of resolution.

The Six Imaging Variables That Decide AI Detection Performance

VariableWhat to MeasureAI-Readiness GuidanceCommon Failure Mode
Pixel density on targetHorizontal resolution ÷ scene width at the priority zone, in px/mClear the detection tier (25 px/m) with margin in every priority zone; more density raises confidence on small objects like firearmsCoverage maps measure floor area, not density; the far half of a scene falls below the floor
CompressionCodec, bitrate, and rate-factor settings on the stream the analytics consumeModerate compression is well tolerated in published testing; avoid aggressive presets on analytics streams, especially for night scenesStorage-era bitrate caps silently degrade small, low-contrast target detection
Frame rateDelivered (not configured) frames per second on each feedEnough samples to catch brief events; verify the encoder is not trading per-frame quality to hold bitrateBrief exposures of a threat fall between analyzed frames
Low light / IRNight-time contrast structure, IR illumination reach and evennessEdges and silhouettes must survive at night in priority zones; lighting upgrades are cheap accuracy gainsNoise floods fine detail after dark in exactly the highest-risk hours
Dynamic rangeSubject detail (not silhouette) at entrances and glass-backed scenesWDR-capable cameras or repositioning at every backlit choke pointEvery subject entering through a backlit doorway arrives as a silhouette
Mounting geometryCamera height, tilt angle, and target aspect in priority zonesModerate angles that preserve body and carried-object aspect; avoid steep top-down views for threat detectionVandal-height mounting foreshortens people and shrinks the effective target footprint

Auditing an existing camera estate for AI readiness

An AI-readiness audit is a desk exercise plus a site walk, and a mid-sized facility can complete one in days, not months. The structure that works in practice has four passes.

Pass one: inventory and triage. Pull the camera schedule from the VMS: model, resolution, lens, mounting location, and the stream settings actually delivered to recording. Most estates discover configuration drift in this pass alone, cameras delivering lower resolution or higher compression than anyone remembers configuring. The integration mechanics of connecting an analytics layer to this inventory, ONVIF profiles, RTSP streams, VMS versions, are treated in depth in the IntelliSee briefing on retrofit architecture for AI physical security.

Pass two: priority-zone density math. Define the zones where detection has operational value: entrances, lobbies, parking structures, perimeter gates, loading docks. For each camera covering a priority zone, compute pixel density at the zone's far edge using the resolution-over-scene-width arithmetic above. Score each zone against the DORI ladder. The output is a density map that usually shows a familiar pattern: strong density at doors, weak density across the open areas where pre-incident behavior actually unfolds.

Pass three: condition survey. Walk the priority zones at night and at peak backlight hours. Note IR performance, WDR behavior at entrances, occlusions that have grown or been built since installation, and dirty or misaimed housings. This pass routinely finds that the cheapest fixes, cleaning, re-aiming, adding a light fixture, recover more detection capability than any hardware purchase.

Pass four: remediation tiering. Sort findings into three buckets: configuration fixes (stream settings, compression profiles, re-aiming) that cost hours; environmental fixes (lighting, vegetation, repositioning) that cost days; and hardware gaps (zones where no realistic configuration reaches the detection floor) that enter the capital plan. The audit's purpose is to operationalize the standard, converting an abstract pixel-density requirement into a ranked work order. In most estates the deployable surface is much larger than the replace-first instinct assumes, which reshapes both the budget conversation and the timeline. Where the inference itself should run, on the camera, on premises, or in the cloud, is a separate architectural decision with its own tradeoffs, analyzed in the companion briefing on edge versus cloud AI inference.

Where retrofit AI fits: working with the cameras you already own

The economic logic of AI video analytics depends on the installed base. Industry analyses place the global surveillance installed base above one billion cameras, and the overwhelming majority were purchased as recording instruments, specified for forensic review rather than real-time machine analysis. A procurement doctrine that requires replacing that infrastructure before analytics can begin prices most organizations out of the category. A doctrine that measures the existing estate against task-anchored density tiers, fixes what configuration can fix, and deploys analytics where the math already works gets organizations into real-time detection at a fraction of the capital cost.

This is the deployment model IntelliSee was built around. The platform connects to existing IP cameras through the facility's VMS, runs detection on an on-premises appliance, and layers drawn-firearm detection, fall detection, and related risk modules onto feeds the organization already owns, with alerts delivered to designated responders within seconds. Because detection operates at the object-signature level rather than the identification tier, it does not require facial recognition, does not store video, and does not collect PHI, and it tolerates the pixel densities that real-world camera estates actually deliver, with per-detection confidence scores that let security teams calibrate alerting to each camera's envelope. The platform's detection-to-alert pipeline and its fit across K-12, healthcare, higher education, manufacturing, and other sectors are documented across the IntelliSee site, and a structured conversation with the IntelliSee team can turn a camera-readiness audit into a deployment plan.

The buyer takeaway from sixty-plus years of imaging science is compact. Detection performance is set upstream of the model, by pixels on target, by compression, by light, and by geometry. Vendors who start the conversation with your camera math are taking the problem seriously. Vendors who promise uniform accuracy across any estate, sight unseen, are not.

Frequently asked questions about camera requirements for AI video analytics

What pixel density does AI threat detection actually need?

As a planning baseline, clear the IEC 62676-4 detection tier of 25 pixels per meter with comfortable margin in every priority zone, and treat higher density as direct fuel for confidence on small targets like drawn firearms. Person-level detections (a person in a restricted zone, a person down) perform well at modest densities; small-object detections benefit meaningfully from every additional pixel on target. The honest answer is per-camera, which is why a readiness audit computes density zone by zone rather than quoting one number for an estate.

Do I need to replace 1080p cameras with 4K for AI video analytics?

Usually not. A 1080p camera with an appropriate lens watching a defined zone often delivers more pixels on target than a 4K camera asked to cover an entire parking lot. Resolution upgrades matter only where the density math fails and no configuration change can fix it. Most estates find their deployable surface is far larger than expected once they measure density instead of megapixels.

How do I calculate pixels per meter for an existing camera?

Divide the camera's horizontal resolution by the width of the scene at the distance that matters. Scene width is twice the distance multiplied by the tangent of half the lens's horizontal field of view. A 1,920-pixel camera with an 80 degree lens covers about 25 meters of width at 15 meters of distance, delivering roughly 76 px/m at that depth. Most VMS planning tools and lens calculators automate this arithmetic.

Does video compression reduce AI detection accuracy?

At moderate settings, published research shows detection is robust: testing across compression levels found constant-rate-factor settings around 37 preserved performance while sharply reducing bitrate. Accuracy degrades at aggressive compression, and the degradation concentrates on small, low-contrast, and fast-moving targets in poorly lit scenes, which is precisely the security-relevant case. Audit the stream the analytics actually consume, not the camera's native capability.

Can AI analytics run on older analog cameras?

Analog feeds digitized through encoders can technically reach an analytics platform, but they typically fail the density and detail math: low effective resolution, interlacing artifacts, and heavy noise leave too little structure for reliable small-object detection. Analog zones are usually the legitimate hardware-replacement line items in a readiness audit, while the IP estate is largely deployable as-is.

Does AI threat detection require identification-tier image quality or facial recognition?

No. The DORI identification tier (250 px/m) exists for establishing who a person is, which is an evidentiary and forensic function. Object-level threat detection establishes what is present and what is happening, and platforms like IntelliSee perform no facial recognition, store no video, and collect no PHI. Operating at the detection and observation tiers is both the privacy posture and the reason retrofit deployments are economically viable.

Continue the research

This briefing covers the imaging chain that feeds the model. For adjacent layers of the architecture:

Request a Camera-Readiness Assessment

Talk to an IntelliSee security specialist. No sales pitch — a structured conversation about your environment, your threat profile, and whether computer vision is the right fit.

Request a Risk Assessment