July 28, 2026 · 8 min read
AI passes the demo. The eighth try is the problem.
What the measurements say about AI in property operations. Some of it is better than you would guess. The phone is a good deal worse, and the way it fails is not the way people warn you it will.
Take one customer service task. Run it through a leading AI eight times in a row. On a published benchmark, the odds that all eight come back correct are under one in four.
The demo you sat through was one run.
That distance between one run and eight is the most useful thing to understand about this technology, and nobody selling it is going to raise it with you. So here is what the research shows about where these systems are genuinely good, where they quietly go backwards, and why the phone is the hardest surface of all. Every number below is tied to a study you can open.
The measured uplift lands on your weakest person, not your best
The best study in this area followed 5,172 customer support agents through a staggered rollout of an AI assistant, across three million chats. The published version in the Quarterly Journal of Economics reports a 15% increase in issues resolved per hour on average.
The average hides the finding that matters. Novice and lower skilled agents improved a lot, by around a third in the working paper version. The most experienced agents saw almost nothing, and the authors report small but statistically significant declines in resolution rates and customer satisfaction among the top performers.
Less experienced and lower-skilled workers improve both the speed and quality of their output, while the most experienced and highest-skilled workers see small gains in speed and small declines in quality.
Brynjolfsson, Li and Raymond, Generative AI at Work, Quarterly Journal of Economics
For an owner, that changes the question. It raises the floor and leaves the ceiling roughly where it was. If your problem is that your best person is your bottleneck, this does less for you than the brochure suggests. If your problem is that a new hire takes nine months to answer a tenant properly, it does more.
One honest caveat. That study ran between 2020 and 2021, on a pre-ChatGPT assistant. The tools are far better now. The pattern of who benefits has held up in later work, but the exact percentage should not be treated as a 2026 number.
It is excellent at the task sitting right next to the one it fails
The clearest description of how AI fails comes from a field experiment with 758 consultants at Boston Consulting Group, published in Organization Science.
On eighteen realistic tasks chosen to sit inside the technology's capability, consultants using AI completed 12.2% more tasks, 25.1% faster, at higher quality. On one complex judgment task chosen to sit just outside it, consultants using AI were 19% less likely to reach the correct answer than the ones working without it.
The researchers call this the jagged frontier. A boundary is not the surprising part. The surprising part is that you can't see where it runs. The task it handles beautifully and the task it gets confidently wrong look identical from the outside.
In a property management company, the tasks sitting just past that line are the ones you would guess. A tenant whose story keeps shifting. An owner conversation about money. Anything touching fair housing, a reasonable accommodation request, or an eviction. Those are complex judgment calls, and complex judgment is exactly where the measured performance goes backwards.
Voice agents complete 31 to 51% of the tasks text handles at 85%
This is the least understood number in the category, and the most important one if somebody is selling you an AI receptionist.
In March 2026 researchers published a benchmark that runs the same customer service tasks through voice agents rather than text. Across 278 tasks, a strong text model completed 85%. Voice agents managed 31% to 51% under clean conditions, and 26% to 38% once you add background noise and a range of accents. The paper's summary: voice keeps only 30% to 45% of text capability. Between 79% and 90% of the failures come from the agent's own behavior rather than from the task.
That is the honest number sitting under every AI phone answering pitch. Speech recognition errors, accents, holding authentication across turns, a dog barking in the background. All of it takes a cut before the model ever reaches the question.
It is also moving quickly, and saying so is part of being straight with you. On the public leaderboard the best voice model went from around 30% in August 2025 to 67% in April 2026. If somebody shows you a demo that works, it probably does work, on a clean line, on the calls it was demonstrated with.
The same task, eight times: leading models score under 25%
A benchmark for customer service agents called tau-bench measures the thing no demo will ever show you. It runs the same task eight times and asks how often the agent gets it right every single time.
On retail tasks, leading models scored under 25% on that measure, while succeeding on individual attempts far more often. Ask the same question eight times and you get eight correct answers less than a quarter of the time.
This is the part worth worrying about, and it is not the part people worry about. A tool that fails obviously is manageable. A tool that answers the same tenant question three different ways depending on the hour is a compliance problem forming quietly, and it will pass every demo you sit through.
Purpose built legal AI still hallucinates 17 to 33% of the time
Stanford researchers ran a preregistered test of the legal research tools sold by LexisNexis and Thomson Reuters, the ones marketed specifically as hallucination free. Across more than 200 legal queries, they hallucinated between 17% and 33% of the time depending on the product. General purpose ChatGPT-4, asked specific verifiable questions about federal court cases, was wrong 58% of the time.
The researchers separate two kinds of error, and the second is the dangerous one. An incorrect answer states the wrong law. A misgrounded answer states the right law and cites it to a source that does not support it. The first kind gets caught. The second reads as authoritative and sails straight past a busy manager.
Now move that to your business. If the expensive, retrieval grounded, purpose built legal product is wrong one time in six, a general chatbot answering a question about your state's notice period isn't something you leave unsupervised.
79% of adults strongly prefer a human, and your owners most of all
A survey of 2,017 US adults in December 2025, with a stated margin of error of 2.5 points, found 79% strongly prefer dealing with a human over an AI agent. Eight percent prefer AI.
The generational split runs the opposite way to the usual assumption about who is nervous here. Fourteen percent of Gen Z respondents preferred AI. Among Boomers it was four percent. Your owner clients are mostly in that last group.
None of which argues against using it. It argues for being deliberate about where it speaks in your voice, and for keeping the route to a person short and obvious.
Under 20% of small US businesses are using AI at all
The US Census Bureau runs a nationally representative survey of business AI use. Between December 2025 and May 2026, 17% to 20% of US businesses reported using AI. Among firms with fewer than 20 employees it was under 20%, and it had not moved significantly across those six months. Among firms with 250 or more employees it was 37%.
So the gap here is about company size rather than about how late you are. If you run a ten person shop and have done nothing with AI yet, you're with the large majority of American businesses, and most of the pressure you feel is coming from people who want to sell you something.
Take the volume, escalate the judgment
Human in the loop gets used as a marketing phrase. The research above is the reason it is not one.
Draft and log, do not send. Take the volume, escalate the judgment. Keep a person on anything involving screening, accommodation requests, pricing, or a safety call. That is what the measurements say, rather than a philosophy we turned up with. High volume, low ambiguity, well defined tasks land inside the frontier. Everything with a judgment call in it lands outside.
The other conclusion is duller and matters more. Don't take anyone's word for whether it's working, including your own team's. Researchers at METR ran a controlled trial in 2025 where experienced developers using AI were measured as 19% slower, and still believed afterwards that they had been 20% faster. A forty point gap between what people felt and what the stopwatch said.
METR has since said that study had a selection problem and is redesigning it, so treat the specific number as unsettled. The gap between perception and measurement is the part that's held up. Whatever you put in, instrument it, and judge it on what changed in your own numbers.
None of this means the technology is stuck
On OSWorld, a benchmark for agents operating a real computer, the best model scored 12.24% in 2024. By July 2026 the leaderboard sat around 85%. Voice more than doubled in eight months.
So the credible position is this. AI can do specific, bounded, well defined parts of this work today, it is improving faster than most software ever has, and it's still nowhere near unsupervised on anything requiring judgment. Buy against the first half of that sentence. Plan for the second.
Sources
- Generative AI at Work5,172 customer support agents, three million chats, staggered rollout from November 2020 to February 2021. Figures here are the published pair: 5,172 agents and 15%. The earlier NBER working paper counted 5,179 agents and reported 14%. The assistant studied predates ChatGPT.
- Navigating the Jagged Technological FrontierField experiment with 758 Boston Consulting Group consultants across 18 tasks inside the capability frontier and one complex managerial task outside it.
- tau-voice: benchmarking voice agents on real customer service tasks278 tasks. Reports pass@1 completion for voice agents against text baselines, under clean conditions and with noise and diverse accents.
- tau-bench: a benchmark for tool-agent-user interaction in real-world domainsIntroduces the pass^8 measure: how often an agent completes the same task correctly on eight consecutive attempts. Leading models scored under 25% on retail tasks.
- Hallucination-Free? Assessing the Reliability of Leading AI Legal Research ToolsPreregistered test of more than 200 legal queries against LexisNexis and Thomson Reuters products. Lexis+ 17 to 18%, Westlaw about 33%.
- Large Legal Fictions: Profiling Legal Hallucinations in Large Language ModelsSpecific, verifiable questions about random federal court cases. Hallucination rates ran from 58% with ChatGPT-4 to 88% with Llama 2.
- Public voice and text agent standingsLive leaderboard, checked 28 July 2026. The best voice model moved from around 30% in August 2025 to 67% in April 2026. Standings change; the figures here are a reading on one date.
- Customer service statistics2,017 US adults, fielded 10 to 11 December 2025, margin of error 2.5 points, weighted for age, race, sex, education and geography. Published by a survey vendor, but the methodology is fully disclosed.
- Business Trends and Outlook Survey: AI use among US businessesNationally representative, biweekly. Reference period 14 December 2025 to 3 May 2026. Real estate and rental and leasing is not broken out separately in this release.
- Measuring the impact of early-2025 AI on experienced open-source developer productivity16 developers, 246 issues. METR later reported a selection bias problem with this study and is redesigning it. We cite the perception gap, which has held up, rather than the headline slowdown figure.
Want the same math run on your portfolio?
We do a 30-minute walkthrough: your door count, your team, where your hours actually go. You leave with the numbers whether or not you work with us.