All articles

July 28, 2026 · 8 min read

AI passes the demo. The eighth try is the problem.

What the measurements say about AI in property operations. Some of it is better than you would guess. The phone is a good deal worse, and the way it fails is not the way people warn you it will.

Written byParcWorkProperty Operations team

Take one customer service task. Run it through a leading AI eight times in a row. On a published benchmark, the odds that all eight come back correct are under one in four.

The demo you sat through was one run.

That distance between one run and eight is the most useful thing to understand about this technology, and nobody selling it is going to raise it with you. So here is what the research shows about where these systems are genuinely good, where they quietly go backwards, and why the phone is the hardest surface of all. Every number below is tied to a study you can open.

The measured uplift lands on your weakest person, not your best

The best study in this area followed 5,172 customer support agents through a staggered rollout of an AI assistant, across three million chats. The published version in the Quarterly Journal of Economics reports a 15% increase in issues resolved per hour on average.

The average hides the finding that matters. Novice and lower skilled agents improved a lot, by around a third in the working paper version. The most experienced agents saw almost nothing, and the authors report small but statistically significant declines in resolution rates and customer satisfaction among the top performers.

Less experienced and lower-skilled workers improve both the speed and quality of their output, while the most experienced and highest-skilled workers see small gains in speed and small declines in quality.

Brynjolfsson, Li and Raymond, Generative AI at Work, Quarterly Journal of Economics

For an owner, that changes the question. It raises the floor and leaves the ceiling roughly where it was. If your problem is that your best person is your bottleneck, this does less for you than the brochure suggests. If your problem is that a new hire takes nine months to answer a tenant properly, it does more.

One honest caveat. That study ran between 2020 and 2021, on a pre-ChatGPT assistant. The tools are far better now. The pattern of who benefits has held up in later work, but the exact percentage should not be treated as a 2026 number.

It is excellent at the task sitting right next to the one it fails

The clearest description of how AI fails comes from a field experiment with 758 consultants at Boston Consulting Group, published in Organization Science.

On eighteen realistic tasks chosen to sit inside the technology's capability, consultants using AI completed 12.2% more tasks, 25.1% faster, at higher quality. On one complex judgment task chosen to sit just outside it, consultants using AI were 19% less likely to reach the correct answer than the ones working without it.

The researchers call this the jagged frontier. A boundary is not the surprising part. The surprising part is that you can't see where it runs. The task it handles beautifully and the task it gets confidently wrong look identical from the outside.

In a property management company, the tasks sitting just past that line are the ones you would guess. A tenant whose story keeps shifting. An owner conversation about money. Anything touching fair housing, a reasonable accommodation request, or an eviction. Those are complex judgment calls, and complex judgment is exactly where the measured performance goes backwards.

Voice agents complete 31 to 51% of the tasks text handles at 85%

This is the least understood number in the category, and the most important one if somebody is selling you an AI receptionist.

In March 2026 researchers published a benchmark that runs the same customer service tasks through voice agents rather than text. Across 278 tasks, a strong text model completed 85%. Voice agents managed 31% to 51% under clean conditions, and 26% to 38% once you add background noise and a range of accents. The paper's summary: voice keeps only 30% to 45% of text capability. Between 79% and 90% of the failures come from the agent's own behavior rather than from the task.

31 to 51%
Share of real customer service tasks completed by voice agents under clean conditions. The same tasks in text run around 85%.Ray et al., tau-voice benchmark, March 2026, 278 tasks.

That is the honest number sitting under every AI phone answering pitch. Speech recognition errors, accents, holding authentication across turns, a dog barking in the background. All of it takes a cut before the model ever reaches the question.

It is also moving quickly, and saying so is part of being straight with you. On the public leaderboard the best voice model went from around 30% in August 2025 to 67% in April 2026. If somebody shows you a demo that works, it probably does work, on a clean line, on the calls it was demonstrated with.

The same task, eight times: leading models score under 25%

A benchmark for customer service agents called tau-bench measures the thing no demo will ever show you. It runs the same task eight times and asks how often the agent gets it right every single time.

On retail tasks, leading models scored under 25% on that measure, while succeeding on individual attempts far more often. Ask the same question eight times and you get eight correct answers less than a quarter of the time.

This is the part worth worrying about, and it is not the part people worry about. A tool that fails obviously is manageable. A tool that answers the same tenant question three different ways depending on the hour is a compliance problem forming quietly, and it will pass every demo you sit through.

Stanford researchers ran a preregistered test of the legal research tools sold by LexisNexis and Thomson Reuters, the ones marketed specifically as hallucination free. Across more than 200 legal queries, they hallucinated between 17% and 33% of the time depending on the product. General purpose ChatGPT-4, asked specific verifiable questions about federal court cases, was wrong 58% of the time.

The researchers separate two kinds of error, and the second is the dangerous one. An incorrect answer states the wrong law. A misgrounded answer states the right law and cites it to a source that does not support it. The first kind gets caught. The second reads as authoritative and sails straight past a busy manager.

Now move that to your business. If the expensive, retrieval grounded, purpose built legal product is wrong one time in six, a general chatbot answering a question about your state's notice period isn't something you leave unsupervised.

79% of adults strongly prefer a human, and your owners most of all

A survey of 2,017 US adults in December 2025, with a stated margin of error of 2.5 points, found 79% strongly prefer dealing with a human over an AI agent. Eight percent prefer AI.

The generational split runs the opposite way to the usual assumption about who is nervous here. Fourteen percent of Gen Z respondents preferred AI. Among Boomers it was four percent. Your owner clients are mostly in that last group.

None of which argues against using it. It argues for being deliberate about where it speaks in your voice, and for keeping the route to a person short and obvious.

Under 20% of small US businesses are using AI at all

The US Census Bureau runs a nationally representative survey of business AI use. Between December 2025 and May 2026, 17% to 20% of US businesses reported using AI. Among firms with fewer than 20 employees it was under 20%, and it had not moved significantly across those six months. Among firms with 250 or more employees it was 37%.

So the gap here is about company size rather than about how late you are. If you run a ten person shop and have done nothing with AI yet, you're with the large majority of American businesses, and most of the pressure you feel is coming from people who want to sell you something.

Take the volume, escalate the judgment

Human in the loop gets used as a marketing phrase. The research above is the reason it is not one.

Draft and log, do not send. Take the volume, escalate the judgment. Keep a person on anything involving screening, accommodation requests, pricing, or a safety call. That is what the measurements say, rather than a philosophy we turned up with. High volume, low ambiguity, well defined tasks land inside the frontier. Everything with a judgment call in it lands outside.

The other conclusion is duller and matters more. Don't take anyone's word for whether it's working, including your own team's. Researchers at METR ran a controlled trial in 2025 where experienced developers using AI were measured as 19% slower, and still believed afterwards that they had been 20% faster. A forty point gap between what people felt and what the stopwatch said.

METR has since said that study had a selection problem and is redesigning it, so treat the specific number as unsettled. The gap between perception and measurement is the part that's held up. Whatever you put in, instrument it, and judge it on what changed in your own numbers.

None of this means the technology is stuck

On OSWorld, a benchmark for agents operating a real computer, the best model scored 12.24% in 2024. By July 2026 the leaderboard sat around 85%. Voice more than doubled in eight months.

So the credible position is this. AI can do specific, bounded, well defined parts of this work today, it is improving faster than most software ever has, and it's still nowhere near unsupervised on anything requiring judgment. Buy against the first half of that sentence. Plan for the second.

Sources

  1. Generative AI at WorkBrynjolfsson, Li and Raymond, Quarterly Journal of Economics · 20255,172 customer support agents, three million chats, staggered rollout from November 2020 to February 2021. Figures here are the published pair: 5,172 agents and 15%. The earlier NBER working paper counted 5,179 agents and reported 14%. The assistant studied predates ChatGPT.
  2. Navigating the Jagged Technological FrontierDell'Acqua et al., Organization Science · 2026Field experiment with 758 Boston Consulting Group consultants across 18 tasks inside the capability frontier and one complex managerial task outside it.
  3. tau-voice: benchmarking voice agents on real customer service tasksRay, Dhandhania, Barres and Narasimhan · 2026278 tasks. Reports pass@1 completion for voice agents against text baselines, under clean conditions and with noise and diverse accents.
  4. tau-bench: a benchmark for tool-agent-user interaction in real-world domainsYao, Shinn, Razavi and Narasimhan · 2024Introduces the pass^8 measure: how often an agent completes the same task correctly on eight consecutive attempts. Leading models scored under 25% on retail tasks.
  5. Hallucination-Free? Assessing the Reliability of Leading AI Legal Research ToolsMagesh et al., Journal of Empirical Legal Studies · 2025Preregistered test of more than 200 legal queries against LexisNexis and Thomson Reuters products. Lexis+ 17 to 18%, Westlaw about 33%.
  6. Large Legal Fictions: Profiling Legal Hallucinations in Large Language ModelsDahl, Magesh, Suzgun and Ho, Journal of Legal Analysis · 2024Specific, verifiable questions about random federal court cases. Hallucination rates ran from 58% with ChatGPT-4 to 88% with Llama 2.
  7. Public voice and text agent standingstau-bench leaderboard · 2026Live leaderboard, checked 28 July 2026. The best voice model moved from around 30% in August 2025 to 67% in April 2026. Standings change; the figures here are a reading on one date.
  8. Customer service statisticsSurveyMonkey · 20252,017 US adults, fielded 10 to 11 December 2025, margin of error 2.5 points, weighted for age, race, sex, education and geography. Published by a survey vendor, but the methodology is fully disclosed.
  9. Business Trends and Outlook Survey: AI use among US businessesUS Census Bureau · 2026Nationally representative, biweekly. Reference period 14 December 2025 to 3 May 2026. Real estate and rental and leasing is not broken out separately in this release.
  10. Measuring the impact of early-2025 AI on experienced open-source developer productivityMETR · 202516 developers, 246 issues. METR later reported a selection bias problem with this study and is redesigning it. We cite the perception gap, which has held up, rather than the headline slowdown figure.

Want the same math run on your portfolio?

We do a 30-minute walkthrough: your door count, your team, where your hours actually go. You leave with the numbers whether or not you work with us.

See all 8 articles