Research

The evidence file.

We keep receipts. This page collects the documented, published deployments of production AI we cite across our work library — who deployed what, what it measurably did, and where the claim was published. None of these are our clients; all of them are checkable. If a class of system appears here, there's a blueprint for how we build it.

88% of organizations now use AI in at least one business function — but only about a third have begun to scale it McKinsey State of AI, 2025 ↗ 95% of enterprise GenAI pilots deliver little to no measurable P&L impact MIT NANDA via Fortune, 2025 ↗ 42% of companies abandoned most of their AI initiatives in 2025 — scrapping an average 46% of proofs-of-concept before production S&P Global, 2025 ↗ $200B in net-new tech-services demand from integrating AI agents into real enterprise systems, over five years BCG, 2026 ↗ 51% of B2B software buyers now begin purchase research in an AI chatbot rather than a search engine — up from 29% a year earlier G2, 2026 ↗ >40% of agentic-AI projects predicted to be canceled by end of 2027 on cost, unclear value, or inadequate risk controls Gartner, 2025 ↗

Yes, we publish the failure numbers. The gap between pilot and production is the entire reason an engineering lab exists — evals, integration, and operations are what separate the 95% from the deployments below.

Kittle Property Group (multifamily operator)

2026 · built with EliseAI

AI leasing agent (LeasingAI) handling prospect inquiries, follow-up, tour scheduling, and delinquency outreach 24/7 across the portfolio.

  • 91% of leasing messages handled by the AI
  • 65% reduction in lead-to-lease timeline
  • 15% increase in tour-to-lease conversion
  • 40% reduction in ad spend
Source: EliseAI customer story ↗

Zillow

2021 · built with in-house

Neural-network Zestimate — rebuilt automated valuation model (AVM) replacing ~1,000 local statistical models with a single national neural network trained on sales transactions, tax assessments, and public records.

  • National median error rate for off-market homes improved to 6.9% (nearly a full percentage point gain)
  • Valuations computed for 104+ million U.S. homes
  • Homes eligible for Zillow Offers cash offers expected to increase ~30% (from ~900,000 homes across 23 markets)
Source: PR Newswire — Zillow press release, June 15, 2021 ↗

AppFolio

2025 · built with in-house

Realm-X AI agents for property management: a Leasing Performer that answers inquiries, manages prospect info and schedules showings, and a Maintenance Performer that diagnoses and prioritizes maintenance requests (including image-based issue detection) and creates work orders.

  • 10 hours per week saved on task management (user surveys, Sep 2024–Feb 2025)
  • 73% higher lead-to-showing conversion rate for Realm-X Flows users vs non-users
  • 26 seconds saved per resident message with Realm-X Messages
Source: AppFolio newsroom announcement, June 11, 2025 ↗

HelloData.ai

2026 · built with in-house

Automated multifamily rent-comp and market-intelligence platform that scrapes property websites and listing sites daily and uses computer vision on listing photos to score unit quality and comparability.

  • 35M+ units surveyed daily with 331M daily prices surfaced
  • 5M+ properties tracked across 99% of U.S. markets
  • 250K+ property websites monitored plus 15+ other public data sources
  • 148M+ listing images processed; 391K concessions analyzed
Source: HelloData.ai — Our Data page ↗

Klarna

2024 · built with OpenAI

Customer-service AI assistant powered by OpenAI, live in the Klarna app, handling multilingual support errands (refunds, returns, payments, cancellations) across 23 markets 24/7.

  • 2.3 million conversations in its first month — two-thirds of Klarna's customer service chats
  • Doing the equivalent work of 700 full-time agents
  • Errand resolution time cut from 11 minutes to under 2 minutes
  • 25% drop in repeat inquiries (greater accuracy)
Source: Klarna press release ↗

Verizon

2025 · built with Google (Gemini, via Google Cloud)

Internal AI assistant for ~28,000 customer service reps, built on Google's Gemini model fed with ~15,000 internal Verizon documents; surfaces answers during calls to cut handling time and free reps to sell. Rolled out July 2024, full scale January 2025.

  • Sales through the 28,000-person service team up nearly 40% since deployment
  • Cut call-handling times, shifting reps from care to revenue-generating work ('reskilling in real time')
  • Built by feeding ~15,000 internal documents into Google's Gemini LLM
Source: Reuters ↗

Air India

2025 · built with Microsoft (Azure OpenAI)

AI.g, a generative-AI virtual assistant on Azure OpenAI with RAG and deep back-end integration, handling booking changes, refunds, and other customer queries as the airline's first line of support.

  • Handles about 40,000 customer queries daily across 1,300+ question types
  • Resolved more than 13 million conversations with a 97% success rate
  • 97% of 4 million-plus customer queries handled with full automation; only ~3% escalated to human agents
  • 'We've saved millions' in support costs; call volumes held flat despite doubling passenger traffic
Source: Microsoft Customer Stories ↗

Octopus Energy

2026 · built with in-house

Arlo, a generative-AI email assistant built on the Kraken platform that autonomously answers routine customer emails (tariff renewals, payment dates, account details), with humans handling complex cases and vulnerable customers.

  • Handled around 8,000 customer emails a week during a three-month UK trial (~4% of all customer emails)
  • 76% customer satisfaction score vs 72% for comparable human-advisor responses
Source: Octopus Energy press release ↗

Lemonade

2023 · built with in-house

AI claims agent ('AI Jim') embedded in Lemonade's app that assesses the claim, checks policy conditions, runs dozens of anti-fraud algorithms, and triggers payment instructions — end to end without a human for eligible claims.

  • Settled a genuine insurance claim (a stolen-bike theft claim in London) in 2 seconds, a claimed world record
  • Nearly half of all Lemonade claims are handled using AI; per Lemonade ~40% require no human intervention at all
Source: Reinsurance News ↗

Zurich Insurance

2017

AI/robotics system that reviews paperwork such as medical reports to decide personal injury claims, deployed from March 2017.

  • Saved 40,000 work hours
  • Cut claim processing time from one hour to five seconds
  • Chairman Tom de Swaan also reported improved accuracy, with machine learning improving on every new claim
Source: Insurance Journal ↗

Fukoku Mutual Life Insurance (Japan)

2017 · built with IBM (Watson Explorer)

IBM Watson Explorer deployment that reads medical documents and calculates policy payouts for claims, replacing manual claims-assessment work (humans still approve final payment).

  • Replaced 34 claims-assessment employees
  • 30% productivity increase in the claims process
  • System cost $1.73M with ~$129K/year maintenance, against expected savings of ~$1.21M per year (ROI within two years)
Source: Fortune ↗

Hiscox (London Market)

2023-2024 · built with in-house

First generative-AI-enhanced lead underwriting model in the London insurance market: Hiscox AI Laboratories (Hailo) plus Google Cloud Gemini/Vertex AI extracts data from broker email submissions and auto-generates a priced quote email ready for underwriter review. Piloted in the sabotage & terrorism line; went live August 2024 with first risk written via broker WTW.

  • Quote turnaround for lead open-market business cut from a manual process that 'can typically take up to three days' to quoting 'within just three minutes'
  • Live in production (not just PoC) as of August 2024, with the first AI-underwritten risk placed through WTW
Source: Hiscox Group newsroom ↗

HSBC

2023 · built with in-house

Google Cloud AML AI — machine-learning transaction monitoring that replaced HSBC's rules-based anti-money-laundering screening, deployed in the UK and Hong Kong and screening over 1.2 billion transactions per month.

  • Identifies 2 to 4 times as much suspicious activity as the previous rules-based system
  • 60% reduction in alert volumes (fewer false positives to investigate)
  • ~2x more financial crime identified in commercial banking, almost 4x more in retail banking
  • Suspicious accounts now detected within 8 days of first alert
Source: Google Cloud blog ↗

Revolut

2024 · built with in-house

Proprietary AI-based fraud detection and financial-crime operation — advanced AI algorithms plus biometric tools and cybersecurity measures, run by a 2,500-strong 24/7 financial crime team.

  • Prevented over £475 million of potential fraud against customers in 2023
  • 2,500-strong, 24/7 financial crime team employing AI-based algorithms and biometric tools
  • Report also documents scam-mix analytics, e.g. 58% of all money lost to scams originated on Meta platforms
Source: Revolut Financial Crime and Consumer Security Report 2023 ↗

Standard Bank (Africa's largest bank by assets)

2020 · built with WorkFusion

Intelligent automation of customer onboarding and KYC on the WorkFusion platform — document processing, identity verification, and account-opening workflows automated across the bank; program running since 2016.

  • Account opening time reduced to 5 minutes (WorkFusion's related materials cite a prior baseline of ~20 days)
  • 1 million transactions automated monthly
  • 60% reduction in vehicle and asset finance customer verification time
  • 300% ROI via citizen-developer program certifying 100+ staff
Source: WorkFusion customer case study ↗

JPMorgan Chase

2017 · built with in-house

COIN (Contract Intelligence) — in-house machine-learning software that interprets commercial loan agreements, automatically extracting key attributes that previously required manual review by lawyers and loan officers.

  • Eliminated 360,000 hours of annual work by lawyers and loan officers
  • Reviews ~12,000 new commercial credit agreements per year, in seconds per document
  • Reduced loan-servicing mistakes that were often attributable to human error in interpreting contracts
Source: ABA Journal ↗

Allen & Overy (now A&O Shearman)

2023 · built with Harvey (OpenAI-backed legal AI startup)

Firm-wide deployment of Harvey, a generative-AI legal platform built on OpenAI models, used for contract drafting, legal research, and document analysis across all practice areas — the first Magic Circle firm to roll out generative AI firm-wide.

  • Deployed to 3,500+ lawyers across 43 offices, operating in multiple languages
  • ~40,000 queries submitted by lawyers during the beta trial (November 2022 to February 2023) before firm-wide rollout
Source: A&O Shearman newsroom ↗

LawGeex (controlled benchmark vs. 20 experienced US corporate lawyers)

2018 · built with in-house

AI contract-review platform tested head-to-head against 20 US-trained corporate lawyers (alumni of Goldman Sachs, Cisco, Alston & Bird, K&L Gates) on risk-spotting in five previously unseen NDAs; study advised by academics from Stanford Law and USC. Note: a controlled benchmark study, not a production deployment — label it as such when citing.

  • AI achieved 94% accuracy at surfacing risks in NDAs vs. 85% average for experienced lawyers (lowest-performing lawyer: 67%)
  • AI completed the review in 26 seconds vs. lawyers' average of 92 minutes (range 51-156 minutes)
Source: Artificial Lawyer ↗

BT (engagement run by Deloitte Legal)

2023 · built with Luminance (deployed by Deloitte)

Deloitte used Luminance's AI contract-analysis platform to analyze and standardize BT's contract estate — 4,500 documents across 14 European entities — for BT's in-house legal team.

  • 4,500 documents across 14 European entities analyzed and standardized in two weeks
  • Deloitte estimates a 50% time savings compared to a manual review
Source: Legal Dive ↗

The Permanente Medical Group (Kaiser Permanente, Northern California)

2025

Enterprise-wide ambient AI scribe rollout: AI listens to patient visits (with consent) and drafts clinical notes for physician review, deployed across primary and specialty care. Evaluated over a 63-week period, October 2023 through December 2024.

  • 7,260 physicians used the AI scribe across 2,576,627 patient encounters in 63 weeks
  • 15,791 hours of documentation time saved (equivalent to 1,794 eight-hour workdays)
  • 84% of physicians reported a positive effect on patient communication; 82% said overall work satisfaction improved
  • 47% of patients noticed their doctor spent less time at the computer; 56% reported a positive impact on visit quality
Source: American Medical Association ↗

University of Michigan Health-West

2024 · built with Nuance (Microsoft) - Dragon Ambient eXperience (DAX)

Ambient clinical documentation (Nuance DAX) that converts the natural doctor-patient conversation into a draft clinical note; studied with 83 primary care clinicians using qualitative and quantitative methods.

  • Clinicians using DAX in >60% of visits saw an additional 12 patients per month on average
  • wRVUs increased by 20 per month for high-frequency users
  • Revenue from additional visits covered the cost of DAX with an additional 80% return on investment
  • Drop in burnout (exhaustion and disengagement) comparable to a clinician going from full-time to part-time work
Source: Healthcare Dive ↗

CHRISTUS Health (15,000 physicians, 600+ care centers)

2024 · built with Abridge

Abridge generative-AI clinical documentation deployed after a pilot, with enterprise-wide rollout to all ambulatory clinicians planned; AI drafts notes from recorded clinical conversations inside the EHR workflow.

  • 78% reduction in clinician cognitive load
  • 60% less time on documentation outside of work hours
  • 40% decrease in physician burnout rate since deploying Abridge
  • 41% increase in clinicians giving undivided attention to patients
Source: Abridge company newsroom ↗

Care New England (Rhode Island health system)

2025 · built with Notable Health

Intelligent automation of radiology prior authorizations and notice-of-admission (NOA) workflows: the platform gathers required clinical data, validates payer criteria, and submits authorizations that previously took staff ~15 minutes each with ~10-day scheduling turnaround.

  • 55% reduction in authorization-related write-offs
  • 80% reduction in authorization turnaround time (from a baseline of nearly 10 days from submission to patient scheduling)
  • 2,841 staff hours saved on NOA and prior authorization work
  • $644K projected write-off and cost savings within 12 months
Source: Notable Health customer story ↗

Walmart

2024 · built with in-house

Used multiple large language models in-house to create or improve product-attribute data across its e-commerce catalog, feeding search, in-store inventory location tools for associates, and delivery picking.

  • Used LLMs to create or improve more than 850 million pieces of data across its product catalog
  • Without generative AI, the process would have required 100 times the headcount to complete in the same amount of time (per CEO Doug McMillon, Q2 FY2025 earnings call)
  • Expanded internal generative AI tool access to 75,000 employees
Source: CIO Dive ↗

Wayfair

2025 · built with Google Cloud (Gemini / Vertex AI)

Deployed Google's Gemini models (Vertex AI) to automatically categorize and enrich product attributes across its ~30-million-item catalog, plus Gemini for Workspace internally.

  • 67% reduction in time needed to curate new and updated product listings
  • Saved hundreds of thousands of dollars via catalog automation
  • Improved some conversion rates by 2%
Source: PYMNTS ↗

Instacart

2023 · built with in-house

Rebuilt its real-time item-availability prediction system as a hierarchical three-part ML model (General / Trending / Real-Time) that scores the in-stock probability of hundreds of millions of grocery items, using stratified scoring so only 'head' items are scored in near-real-time.

  • Reduced computation costs by about 80% via stratified scoring schedules
  • Predicts availability for hundreds of millions of items across nearly 100,000 stores in more than 15,000 cities in North America
  • Only ~1% of unique items require real-time scoring; ~85% (torso items) scored on monthly intervals
Source: Instacart engineering blog ↗

OTTO (Germany)

2017–present · built with in-house

AI sales-forecasting and autonomous stock-replenishment system that forecasts SKU-level demand up to 450 days ahead and reorders merchandise automatically without human intervention. (The Economist, Apr 2017, separately reported 90% accuracy on 30-day forecasts and a 20% inventory reduction — verified via a Harvard Business School AI Institute writeup citing it.)

  • Over 30 billion individual demand forecasts per month
  • Around 2.5 million article numbers analyzed/forecast for the next 450 days
  • 35% of product ranges already reordered fully automatically
  • Aims to save tens of millions (EUR) in the current financial year through forecast-based automated management
Source: OTTO corporate newsroom ↗

Google

2016 · built with in-house

Ensemble of deep neural networks trained on historical data-center sensor data to predict PUE, temperature and pressure an hour ahead, and recommend cooling-plant settings; deployed live on Google data centers.

  • Reduced energy used for cooling by up to 40 percent
  • 15 percent reduction in overall PUE overhead after accounting for electrical losses
  • Produced the lowest PUE the site had ever seen
Source: Google DeepMind blog ↗

Siemens (Electronics Works Amberg, Germany)

~2021 · built with in-house

AI model that learns from soldering process data to predict PCB solder-joint quality and decide which boards actually need end-of-line X-ray inspection; trained in the cloud, deployed on an Industrial Edge app on the production line at the Amberg plant (which produces ~17 million Simatic components/year across ~1,200 variants).

  • Avoided a €500,000 capital investment in an additional X-ray inspection machine by predicting solder-joint quality instead of testing every board
  • Closed-loop analytics feed AI results straight back into production optimization at a plant running ~350 changeovers per day
Source: Siemens company insights ↗

PepsiCo (Frito-Lay plants)

2022 · built with Augury

AI-driven machine-health / predictive-maintenance platform: wireless vibration/temperature sensors on production machinery with machine-learning diagnostics that flag developing failures before breakdown, deployed across Frito-Lay snack plants via PepsiCo Labs.

  • Manufacturing capacity rose by 4,000 hours a year across four Frito-Lay plants — 'the equivalent of several million pounds of snacks' (PepsiCo Labs GM Anna Farberov, as reported by the WSJ)
  • Reduced unexpected breakdowns, interruptions, and incremental replacement-part costs on monitored assets (PepsiCo supply chain director)
Source: Augury blog citing WSJ CIO Journal ↗

Audi (Neckarsulm plant, Volkswagen Group)

2023 · built with in-house

AI system for quality control of resistance spot welds in body construction: machine-learning analysis of welding process data flags anomalous welds, replacing manual ultrasound spot-checking; being rolled out to further VW Group plants (Brussels, Emden, Ingolstadt).

  • AI analyzes around 1.5 million spot welds on 300 vehicles per shift at Neckarsulm — vs. the previous manual ultrasound method that sampled about 5,000 spot welds per vehicle
  • Staff now focus only on flagged anomalies, enabling more targeted quality control; infrastructure for three more VW Group plants planned by end of 2023
Source: Audi MediaCenter press release ↗

UPS

2016 · built with in-house

ORION (On-Road Integrated Optimization and Navigation) — an in-house route-optimization system layering predictive algorithms over UPS's package and vehicle tracking data to compute daily delivery routes for its US driver fleet.

  • over $320 million in cumulative cost savings by December 2015
  • $300–$400 million projected annual savings at full deployment
  • 10 million gallons of fuel saved annually
  • 100,000 metric tons of CO2 emissions avoided per year
Source: INFORMS ↗

Walmart

2024 · built with in-house

Route Optimization — AI-powered middle-mile logistics software (automated route mapping, trailer-packing optimization, delivery planning) built and used internally, then launched as a SaaS product via Walmart Commerce Technologies in March 2024.

  • eliminated 30 million unnecessary miles driven
  • avoided 94 million pounds of CO2 emissions
  • optimized routes to bypass 110,000 inefficient paths
  • won the 2023 Franz Edelman Award for deploying the technology at scale
Source: Walmart corporate newsroom ↗

DHL Supply Chain

2025 · built with HappyRobot

Agentic AI communication agents (built with startup HappyRobot) that automate operational phone and email workflows — carrier appointment scheduling, driver follow-up/transport-status calls, and high-priority warehouse coordination.

  • slated to save millions of minutes of phone time per year
  • handles hundreds of thousands of emails annually
Source: Yahoo Finance ↗

Kimberly-Clark

2024 · built with FourKites

FourKites Dynamic Yard — real-time yard visibility and orchestration (digital appointment/dock scheduling and trailer tracking) deployed across Kimberly-Clark distribution sites to cut carrier detention.

  • slashed detention fees by 52% in just 30 days
  • 80% reduction in time spent managing site calendars (companion FourKites case study)
Source: FourKites customer case study ↗

Alaska Airlines

2021 · built with in-house

Flyways AI, a flight-routing platform built by Air Space Intelligence that uses AI/ML over weather, winds, turbulence, airspace constraints and traffic data to recommend optimized, ATC-compliant routes to dispatchers. Alaska was the first airline worldwide to adopt it, signing a multi-year contract after a six-month trial.

  • Saved 480,000 gallons of fuel in the six-month trial
  • Avoided 4,600 tons of carbon emissions
  • Identified route/fuel optimization opportunities on 64% of mainline flights; dispatchers accepted 32% of recommendations
Source: Alaska Airlines Newsroom ↗

American Airlines

2021-2023 · built with in-house

Smart Gating, a machine-learning system built on Microsoft Azure that analyzes real-time flight, routing and runway data to automatically assign arriving aircraft to the nearest available gate, first deployed at Dallas Fort Worth in 2021 and since rolled out to CLT, MIA, DCA and ORD.

  • Reduces total aircraft taxi time by 17 hours per day across its five largest hubs
  • Cut taxi times at DFW by 20% (about two minutes per flight, more than 11 hours of taxiing daily)
  • Saves an estimated 1.4 million gallons of jet fuel annually
  • Reduces CO2 emissions by more than 13,000 metric tons annually
Source: American Airlines Newsroom ↗

Booking.com

2019 · built with in-house

A portfolio of roughly 150 customer-facing machine-learning applications across the booking funnel (recommendations, search ranking, personalization, customer service), each deployed to production and evaluated with randomized controlled trials, documented in a peer-reviewed KDD 2019 paper.

  • About 150 successful customer-facing ML applications in production, built by dozens of teams
  • Exposed to hundreds of millions of users worldwide
  • Every model validated through rigorous randomized controlled trials before rollout
Source: ACM SIGKDD 2019 ↗

Odyssey Resorts (nine-property resort operator)

2025 · built with IDeaS Revenue Solutions (SAS)

IDeaS G3 RMS, an AI-driven hotel revenue management system providing automated demand forecasting and dynamic room pricing, deployed portfolio-wide to replace manual, intuition-based pricing. IDeaS (a SAS company) is the market-leading hotel RMS vendor, also selected by Accor as exclusive global RMS provider across 5,600+ hotels in 2024.

  • 8% portfolio RevPAR growth year-over-year through August 2025
  • Five of nine properties achieved double-digit RevPAR growth; top performer reached 19%
  • August 2025 was the first month since 2021 that all nine properties exceeded prior-year lodging revenue simultaneously
Source: IDeaS client success story ↗

Octopus Energy (Kraken Technologies)

2019-2025 · built with in-house

Kraken, an AI/ML-driven operating system for utilities that automates billing, customer operations and the energy supply chain, and manages residential flexibility assets (EVs, home batteries, heat pumps) for grid balancing; licensed to major utilities including EDF, E.ON Next, National Grid US, Origin Energy, Plenitude and Tokyo Gas.

  • Contracted to manage over 70 million household and business customer accounts worldwide
  • Over $500 million in committed annual recurring revenue, quadrupling contracted revenue in three years
  • Processes 15 billion new data points per day across licensed utility clients
Source: Octopus Energy Group press release ↗

Google (wind farm fleet, central United States)

2019 · built with in-house

DeepMind neural-network forecasting applied to Google's 700 MW wind fleet: trained on weather forecasts and historical turbine data to predict wind power output 36 hours ahead, enabling optimal hour-by-hour delivery commitments to the grid a full day in advance.

  • Boosted the value of the wind energy by roughly 20% versus the baseline of no time-based grid commitments
  • Predicts wind power output 36 hours ahead of actual generation
  • Applied across 700 megawatts of wind capacity in the central US
Source: Google DeepMind blog ↗

Unilever

2019 · built with in-house

HireVue AI-scored video interviewing deployed globally in Unilever's graduate recruitment (first trialled 2017), automatically assessing candidates' recorded interview responses against traits predictive of job success to replace early-round human screening.

  • 100,000 hours of human interviewing/recruitment time saved per year (Unilever spokesperson, on the record)
  • Roughly $1 million in recruitment cost savings each year globally
  • Company reports contribution to a more ethnically and gender-diverse intake (demographic data not fully representative, per Unilever)
Source: The Guardian, 'Unilever saves on recruiters by using AI to assess job interviews' ↗

Chipotle Mexican Grill

2024-2025 · built with in-house

'Ava Cado', a conversational-AI hiring assistant built on Paradox, rolled out across 3,500+ restaurants (announced Oct 2024): chats with candidates in four languages, answers questions, collects information, schedules interviews for hiring managers and sends offers.

  • Time from application to start date cut from 12 days to 4 (75% reduction)
  • Application completion rate rose from 50% to 85%
  • 100% increase in application volume
Source: Paradox case study: 'Reducing time to hire by 75% with conversational AI' ↗
Original research

Our own numbers, published next.

In preparation

What LLM extraction actually costs at web scale

Tier-by-tier economics from a production pipeline processing 4,900+ properties a day: what share resolves deterministically, what the LLM fallback really costs per thousand records, and where models earn their keep. Original data from our own systems — the kind of receipt we wish more vendors published.

Evidence says this class of system works.
We'll tell you what it takes in your stack.