Est.

Monitoring Center Operator-to-Camera Ratio Benchmarks

The industry's staffing benchmark has no source and collapses under real-world conditions.

Senior Writer · · 11 min read
Cover illustration for “Monitoring Center Operator-to-Camera Ratio Benchmarks”
Vendor Evaluation · October 3, 2026 · 11 min read · 2,390 words

The number that governs nearly every conversation about monitoring center staffing is a range: 9 to 16 video feeds per operator, for no more than 20 minutes at a stretch. Security directors cite it, vendors build sales materials around it, and training programs treat it like a settled fact. Nobody, however, can point to the study that produced it. IPVM's practitioner discussions show the same figure circulating through training decks and vendor pitches, but nobody attaches a name to it, no author, no year, no methodology anyone can inspect. A number with no traceable origin cannot be tested against new camera types, revised as technology changes, or adjusted for a warehouse versus a lobby. It just repeats.

The gap between that number and daily practice is wide. Real security operations centers and central stations run operators at 20 to 30 cameras each, roughly double the textbook figure, and they get away with it because the operators are not actually watching continuously. They are handling phone calls, logging reports, and checking in on feeds between other tasks. The ratio was never describing full attention in the first place, only a body present at a console. Once "monitoring" and "checking in periodically" start meaning the same thing on the floor, the number loses whatever authority it had to begin with.

How human attention degrades across a monitoring shift

Even granting the benchmark its best-case assumptions, it only describes the opening minutes of a shift. It doesn't hold steady at whatever level the ratio implies. It drops fast, and it keeps dropping. If an operator watches video continuously for roughly 12 minutes, they miss close to half of what happens on screen. By the 22-minute mark, the miss rate climbs close to total. A ratio that looked defensible when an operator sat down at 7 a.m. describes something close to inattention by the time that same operator reaches hour six.

The decline is not uniform across every kind of scene, either. Spotting someone climbing a fence in an empty lot takes almost no cognitive effort because it is visually loud and sits against a static background. Noticing a propped door, or confirming that every forklift in a warehouse bay is parked where it should be, demands sustained, active attention of a kind that erodes far faster than a glance-and-move-on scan. Operators fix this in practice by slicing tasks across many feeds instead of staring continuously at a few, which trades depth of coverage for length of shift. If an operator manages 48 or more cameras across eight hours at roughly half effectiveness, they have stopped monitoring in any meaningful sense and started sampling. Training programs answer this by pointing to fatigue management and attention rotation, and rotation does help, but only at the margins. It does not change what the attention curve does to a person sitting at a console for hours. No supervisor can reschedule a break in a way that repeals that curve, and a threat that occurs at hour six faces a materially different operator than the one who clocked in at hour zero. Degraded attention is the baseline condition of a shift. What pushes that baseline past its limit is what arrives on top of it: the alarm queue.

Diagram: How Operator Attention Collapses Across a Monitoring Session. Visualizes: Show the steep, non-linear decay of operator attention during continuous video monitoring.

Alarm volume and the ratio's queue problem

A ratio that might hold up when things are quiet collapses once alarm volume climbs, because the operator's queue fills with low-value noise, and that crowds out review of the events that actually matter. The industry's reported false-alarm rate runs overwhelmingly high: the vast majority of alarms arriving at a monitoring center turn out to be nothing, a bird on a sensor, a shadow on a motion detector, a branch in the wind. Each one still takes time to open, assess, and close out.

Salt Lake City's experience with verified response shows just how much that waste adds up to. The police department's own summary reports an immediate 90% reduction in alarm responses once verification requirements went into effect. Most of what had been arriving at dispatch did not warrant a response. Every one of those false alarms, before verification filtered them out, had already consumed operator time just to get evaluated.

The arithmetic that follows is simple and brutal. If nearly all incoming alarms are false, and each one takes attention to process, then a genuine threat arriving during a peak-volume period lands in the same queue as the noise and waits its turn behind it. Operators under that kind of load develop coping habits: muting cameras that generate constant chatter, batch-dismissing whole classes of alarm, quietly deprioritizing accounts that cause the most noise. Each of those habits is a sensible response to an impossible workload, and each one trades verified coverage for throughput. Supervisors sometimes point to informal site knowledge, the idea that operators learn over time which cameras cry wolf and adjust accordingly. That knowledge helps at the edges, but it lives in one person's head, it cannot be audited, and it vanishes the moment that operator leaves the job.

Why scaling headcount does not solve the architectural problem

The obvious answer, hire more operators, preserves the ratio on paper without fixing what the ratio was supposed to measure. The bottleneck was never the headcount, but a review model where every single alarm, real or false, has to pass through a human being in sequence before anyone can act on it. A large site running at a high camera-to-operator ratio would need many operators simultaneously at full attention to actually hold that ratio, a staffing level that no monitoring center commits to for a single account, particularly overnight when fewer people are on the floor to begin with.

The economics do not scale in a straight line either. Operator cost is fixed per person regardless of how many alarms come in on a given night, so the periods when the queue runs longest, nights, weekends, weather events, are exactly the periods when the ratio is least likely to hold. The industry's usual workarounds, offshoring to cut labor cost, muting cameras to cut noise, charging overage fees when alarm volume runs high, each trade one problem for another: lower cost for less oversight, less noise for less coverage. None of them touch the queue itself.

Meanwhile the number of cameras that need watching keeps growing. Mordor Intelligence puts the surveillance camera market at roughly $46.69 billion in 2026, and expects double-digit growth to continue through 2031. Camera counts at monitored properties are climbing faster than any realistic operator headcount can be hired, trained, and retained to track them. If how alarms get reviewed doesn't change, more cameras just mean a worse version of the same queue problem every year. What the structure actually needs is something that can sit in front of that queue, process every alarm without fatigue, and pass along to a human only the ones that have already been verified as real.

What a monitoring architecture built around pre-filtering looks like

That is the function an AI-agent-first architecture is built to perform. Traditional monitoring puts a human operator first in the review sequence, so every alarm, false or real, lands in front of a person before anything gets decided. An AI-first model inverts that order: AI reviews every incoming alarm, classifies it, and escalates to a human only the events that meet verified-threat criteria. The noise stops at the AI layer and never reaches the console.

This shift is event-based, not screen-based. Instead of an operator scanning a wall of static feeds and hoping to catch something, the AI layer watches every camera all the time and brings in a human operator only when an event needs attention. That shift lets one trained operator cover a number of cameras that no live-view ratio could sustain, because the operator is no longer paying the cognitive cost of watching quiet footage. One monitoring center that deployed this kind of pre-filtering reported a large drop in alerts per site, hundreds of thousands fewer each month, so existing staff could manage more accounts while carrying meaningfully less cognitive load per shift.

The AI's job stops at filtering. It does not dispatch, it does not speak to anyone on site, and it does not take action. The human operator still makes every dispatch call, issues every audio warning over the speaker, and carries out every response. What the AI changes is simpler and more consequential: it keeps a real threat from sitting in a queue behind a hundred false alarms. That distinction only holds, though, if the filtering is actually tuned to the site rather than applied generically. An AI agent that knows a property's schedules, its authorized activity, and the specific placement of its cameras can distinguish an employee working late from someone who should not be there. A generic motion detector cannot make that distinction no matter how it is tuned; site calibration, not just AI involvement, separates real filtering from a smarter version of muting a noisy camera.

Response time benchmarks and the gap between ratio compliance and actual coverage

A buyer can hold a vendor to response time to a verified threat, not cameras per operator. A ratio tells a buyer how many feeds a person is assigned. It says nothing about how long an actual intrusion sits before a trained human lays eyes on it, and that lag is what stops or fails to stop it.

AI-augmented systems compress that window sharply. The AI layer flags an event within seconds of it occurring, and the monitoring center's service-level architecture then moves that flagged event through triage fast, so a human operator steps in within a short window of the event starting if it's marked as priority. Monitoring centers quoting commercial deployments as of 2026 are setting very short reaction-time targets that bundle in two-way audio warnings, PTZ camera zoom, siren activation, and dispatch coordination, all within that same short window. A perimeter breach caught in a fraction of a second can trigger a siren or move a guard toward the scene before the intruder gets past the fence line. The same breach caught 20 minutes later produces a police report instead of a prevented incident.

The fragility of policy-based fixes makes the case for building verification into the system itself rather than relying on outside rules. Where cities have required verified response before dispatching police to an alarm, alarm responses have dropped sharply, as Salt Lake City shows. But the Security Industry Association reports that only about 19 of roughly 18,000 U.S. law enforcement agencies have formally adopted verified response, and at least eleven have reversed course after adopting it. Policy-driven verification has largely stalled as an industry-wide solution. Verification that depends on winning over thousands of individual police departments is not going to arrive on any predictable timeline. The monitoring system itself has to do that verification work before an alarm ever reaches a dispatcher. For a buyer evaluating vendors, a service agreement that specifies a camera-to-operator ratio promises nothing about how fast a real threat gets escalated. A service agreement that commits to a time-to-verified-escalation number is the one that actually reflects what coverage a site is getting, and what makes that number achievable in the first place is how well the system has been calibrated to the site it is protecting.

Site-specific protocol coverage and whether the ratio math holds

Even the best-architected pre-filtering system fails without that calibration, because if you apply the same detection threshold across every property type, you get back the false-alarm volume that broke the ratio to begin with. A truck yard at two in the morning with an authorized driver making a pickup looks, to a generic motion sensor, exactly like an intrusion. Telling the two apart depends entirely on whether the system knows the driver's schedule, recognizes the vehicle, and understands what a normal access pattern looks like at that specific gate.

Protocol specificity is what actually governs detection quality at the source. Person-only analytics, tight detection zones drawn around high-value areas instead of blanket motion sensing, and webhooks that only pass along already-verified event classes to the alarm system all cut the inbound alarm load before it ever reaches a queue, rather than relying on an operator to sort it out afterward. Different event types also call for different response speeds: a perimeter breach needs a faster window than a loitering alert at a building entrance, which in turn needs a different window than a PPE violation flagged on an industrial floor. A system with no site calibration treats all three the same way, which is itself a design failure.

Commercial properties, retail stores and employee parking lots among them, benefit from filtering that distinguishes someone walking through a lot during business hours from someone doing the same thing at 3 a.m. The value an AI agent delivers scales directly with how precisely it understands what normal, authorized activity looks like at that particular address. Multifamily housing presents the opposite calibration challenge: lobbies, elevators, and parking garages carry heavy volumes of legitimate foot traffic, so the dominant risk is over-triggering rather than missing something. A system tuned to tell resident activity apart from actual threat activity is what makes a verbal warning over a speaker work as a deterrent rather than training everyone in the building to ignore it. Industrial sites carry a further layer: detecting missing protective equipment or someone entering a restricted zone requires models trained against the specific camera angles and layout of that site. The Port of Virginia's deployment, as reported by Voxel, cut daily manual footage review from several hours down to a fraction of that time after continuous AI monitoring went in. Pricing in the industry reflects that this calibration carries real value: AI-enhanced monitoring commands a meaningfully higher per-camera monthly rate than basic alarm-triggered monitoring, a premium built on reduced false alarms and verified response rather than on camera count alone.

None of this is reason to keep asking how many cameras one person can watch. What actually protects a site is the speed with which a verified threat reaches a trained human and the accuracy with which the system tells threats apart from noise given its own schedules, its own layout, and its own authorized traffic. The ratio measures what gets assigned. It has never measured what gets caught.

Sources

  1. How Many Camera Views Can Be Effectively Monitored? - IPVM Discussions
  2. Surveillance Camera Market Size, Trends & Analysis 2031

More in Vendor Evaluation