• Anthropic’s landmark $1.5B copyright settlement is approved• Trump’s latest AI czar has already resigned• Google is working on a new AI chip designed to make Gemini more efficient• AI’s most important protocol is getting a little bit easier to use• X relaunches a rebuilt Android app after year-long effort• OpenAI is scared of open-weight models. Should the US be?• Adobe camera app’s new feature will critique your photos using AI• YouTube clarifies policies around AI slop and upsetting videos• What to watch for after Jensen Huang’s Japan visit• Can an Apple lawsuit derail OpenAI’s hardware plans?• ‘Odyssey’ director Christopher Nolan calls AI an obvious ‘Trojan horse’• Nonprofit Current AI is racing to build the World Wide Web of AI, free for all• Kimi: Threat or menace?• Neil Rimer thinks the AI money is coming back out• Vertu wants executives to pay $6,880 for an AI agent — here’s how it actually performs• Connect more of your apps to Search• Create, edit and star in videos with two Google Vids updates• Celebrating 25 years of visual search innovation• Expanding Managed Agents in Gemini API: background tasks, remote MCP and more• The latest AI news we announced in June 2026• New York City educators and industry leaders gathered at Google’s offices to shape the future of AI in classrooms.• Unlocking Britain’s next era of productivity: Building a nation of AI trailblazers• Ask an AI expert: What exactly is the full stack?• Our latest Google Finance upgrades, including a new app• New research shows how AMIE, our medical AI, could help manage health conditions.• We’re strengthening our presence in Alabama through new investments and community support.• Our new community investments in Virginia support local jobs and expand energy affordability.• The latest AI news we announced in May 2026• 5 ways Google Search can level up your thrift and vintage shopping• How we used Gemini to build Google I/O 2026• AI is writing, acting and producing China’s minidramas. It’s shaking a $14 billion industry. - NBC News• Election voting advice from AI chatbots ‘inaccurate and unreliable’ - The Guardian• London firms warn of widening AI skills gap - BBC• How Google’s A.I. Search Is Imperiling the Open Web - The New York Times• China weighs tighter export controls on AI models and chips - Financial Times• Opinion | Powerful AI models are being given away for free. It was inevitable. - The Washington Post• Head of US AI safety agency resigns - Reuters• These are the most urgent AI risks, according to 272 experts - MIT Sloan• CNBC's The China Connection newsletter: The AI consumer bet might surprise you - CNBC• OpenAI Appears to Be Missing Its Sales Goals by a Vast Margin - Futurism• Artificial Intelligence - ABC News - Breaking News, Latest News and Videos• Nebraska Snapshot shows AI awareness differs by age, education, income - Nebraska Today• UA Little Rock Professor Nitin Agarwal’s AI Research Published at World's Top Artificial Intelligence Conference - UA Little Rock• Frederick Health Hospital using artificial intelligence to advance medical treatment for patients vulnerable to seizures - DC News Now• USC at the Leading Edge of AI - USC Today• Safety and alignment in an era of long-horizon models• A scorecard for the AI age• Why teens deserve access to safe AI• How Cars24 scales conversations and builds faster with OpenAI• The US is advancing AI safety through state and federal action• GPT-Red: Unlocking Self-Improvement for Robustness• How to manage AI investments in the agentic era• How sales teams use ChatGPT Work• How data science teams use ChatGPT Work• How Deutsche Telekom is rewiring telecommunications with AI• Getting started with ChatGPT• GPT-5.6 is now the preferred model in Microsoft 365 Copilot• GPT-5.6: Frontier intelligence that scales with your ambition• ChatGPT is now a partner for your most ambitious work• GPT-5.5 Bio Bug Bounty• 5 ways to build a side hustle with Gemini• How Gemini is speaking the language of Southeast Asia• Here’s how to make study notebooks in the Gemini app.• 3 ways this coffee shop is growing with Gemini• The latest AI news we announced in June 2026• Gemini Spark updates: macOS launch, connected apps and more• Start building with Nano Banana 2 Lite and Gemini Omni Flash• The Gemini app is bringing personalized image creation to more users.• Gemini can now take notes in Google Meet for Google AI Pro and Ultra subscribers.• Here's how Gemini can help you avoid jetlag.• Try these 3 Google AI tools to help find your next job.• 5 ways Google parents are using Gemini• 5 ways to learn with study notebooks in the Gemini app• Introducing computer use in Gemini 3.5 Flash• Powering the world’s first AI arts museum• The AI compute gap: Enterprises are buying infrastructure faster than they can measure what it costs• The agent security gap: 54% of enterprises have already had an AI agent incident, and most still let agents share credentials• The AI context gap: Enterprise AI organizations have a trust problem, not a retrieval problem — and most are still building the fix• The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway• Agentic orchestration: Enterprise AI organizations have a deployment problem, not a platform problem — and most are calling chatbots agents• Google just redesigned the search box for the first time in 25 years — here’s why it matters more than you think.• Railway secures $100 million to challenge AWS with AI-native cloud infrastructure• Best Universities To Study AI in 2026• 10 top women in AI in 2026• Pope Leo XIV Declares AI a Threat to Human Dignity and Workers’ Rights• ChatGPT Is Making People Think They’re Gods and Their Families Are Terrified• AI May Soon Help You Understand What Your Pet Is Trying to Say• Netflix Adds ChatGPT-Powered AI to Stop You From Scrolling Forever• Murder Victim Speaks from the Grave in Courtroom Through AI• China Unveils World’s First AI Hospital: 14 Virtual Doctors Ready to Treat Thousands Daily• Katy Perry Didn’t Attend the Met Gala, But AI Made Her the Star of the Night• Therapists Too Expensive? Why Thousands of Women Are Spilling Their Deepest Secrets to ChatGPT• Relay.app is shutting down: How to export your workflows and move to Zapier• I've tried every automation software: here are the 10 best in 2026• Web scraping: A comprehensive guide• Turn off your slop cannon• The 6 best vibe coding tools in 2026• OpenClaw vs. Zapier: What's the difference? [2026]• Agentic AI vs. RPA: Everything you need to know• 16 AI prompt templates for better AI agent outputs• The best CRM software for real estate agents in 2026• Integrately vs. Zapier: Which is best? [2026]• Workato vs. Zapier for large businesses: Which is best? [2026]• Zapier vs. Gumloop: Which is best? [2026]• AI agent frameworks: Definition, comparison, and guide• The 4 best read it later apps to save content in 2026• The 8 best data integration tools in 2026
Google is working on a new AI chip designed to make Gemini more efficient
AI News & Artificial Intelligence | TechCrunch

Google is working on a new AI chip designed to make Gemini more efficient

Alphabet, Google's parent company, is reportedly working on a new chip designed to make its Gemini models run much more efficiently.

How to manage AI investments in the agentic era
OpenAI News

How to manage AI investments in the agentic era

Learn how enterprises can manage AI investments in the agentic era by measuring useful work per dollar, improving efficiency, and scaling high-value workflows.

Connect more of your apps to Search
AI

Connect more of your apps to Search

You’ll be able to securely link and interact with your go-to services directly in AI Mode.

Ask an AI expert: What exactly is the full stack?
AI

Ask an AI expert: What exactly is the full stack?

A Google expert explains what it means to take a full-stack approach to AI and why it’s been the foundation of our AI work for so long.

China weighs tighter export controls on AI models and chips - Financial Times
"artificial intelligence" - Google News

China weighs tighter export controls on AI models and chips - Financial Times

China weighs tighter export controls on AI models and chips  Financial Times

A scorecard for the AI age
OpenAI News

A scorecard for the AI age

Sarah Friar, CFO of OpenAI, introduces a practical AI scorecard to measure ROI through useful work, cost per successful task, dependability, and return on compute.

UA Little Rock Professor Nitin Agarwal’s AI Research Published at World's Top Artificial Intelligence Conference - UA Little Rock
"artificial intelligence" - Google News

UA Little Rock Professor Nitin Agarwal’s AI Research Published at World's Top Artificial Intelligence Conference - UA Little Rock

UA Little Rock Professor Nitin Agarwal’s AI Research Published at World's Top Artificial Intelligence Conference  UA Little Rock

CNBC's The China Connection newsletter: The AI consumer bet might surprise you - CNBC
"artificial intelligence" - Google News

CNBC's The China Connection newsletter: The AI consumer bet might surprise you - CNBC

CNBC's The China Connection newsletter: The AI consumer bet might surprise you  CNBC

Opinion | Powerful AI models are being given away for free. It was inevitable. - The Washington Post
"artificial intelligence" - Google News

Opinion | Powerful AI models are being given away for free. It was inevitable. - The Washington Post

Opinion | Powerful AI models are being given away for free. It was inevitable.  The Washington Post

Expanding Managed Agents in Gemini API:  background tasks, remote MCP and more
AI

Expanding Managed Agents in Gemini API: background tasks, remote MCP and more

We’re announcing new capabilities in Managed Agents in Gemini API so developers can build reliable, production-ready agents.

The US is advancing AI safety through state and federal action
OpenAI News

The US is advancing AI safety through state and federal action

OpenAI outlines a “reverse federalism” approach to AI governance, where state laws help build a national framework for safe, democratic AI.

Election voting advice from AI chatbots ‘inaccurate and unreliable’ - The Guardian
"artificial intelligence" - Google News

Election voting advice from AI chatbots ‘inaccurate and unreliable’ - The Guardian

Election voting advice from AI chatbots ‘inaccurate and unreliable’  The Guardian

How we used Gemini to build Google I/O 2026
AI

How we used Gemini to build Google I/O 2026

Learn how Googlers used AI to produce Google I/O 2026.

Adobe camera app’s new feature will critique your photos using AI
AI News & Artificial Intelligence | TechCrunch

Adobe camera app’s new feature will critique your photos using AI

Adobe's Project Indigo can now remove all kinds of backgrounds from photos you snap using the app.

Vertu wants executives to pay $6,880 for an AI agent — here’s how it actually performs
AI News & Artificial Intelligence | TechCrunch

Vertu wants executives to pay $6,880 for an AI agent — here’s how it actually performs

From AI workflows to battery life and security, here's what it's really like to live with Vertu's luxury foldable every day.

New York City educators and industry leaders gathered at Google’s offices to shape the future of AI in classrooms.
AI

New York City educators and industry leaders gathered at Google’s offices to shape the future of AI in classrooms.

Google, the New York Jobs CEO Council and Urban Assembly hosted an AI summit for 150 education and industry leaders.

Our latest Google Finance upgrades, including a new app
AI

Our latest Google Finance upgrades, including a new app

The new Google Finance is coming out of beta and launching a new Android app.

Celebrating 25 years of visual search innovation
AI

Celebrating 25 years of visual search innovation

Google Images is turning 25. Here’s a look back at some major milestones — and new ways to explore and create visual content.

Create, edit and star in videos with two Google Vids updates
AI

Create, edit and star in videos with two Google Vids updates

Gemini Omni and personal avatars in Google Vids make video creation easier than ever.

Why teens deserve access to safe AI
OpenAI News

Why teens deserve access to safe AI

Learn how OpenAI is making ChatGPT safer for teens with age-appropriate protections, learning tools, parental controls, and expert partnerships.

New research shows how AMIE, our medical AI, could help manage health conditions.
AI

New research shows how AMIE, our medical AI, could help manage health conditions.

Research in “Nature” shows our conversational AI system matches primary care physicians in complex disease management.

What to watch for after Jensen Huang’s Japan visit
AI News & Artificial Intelligence | TechCrunch

What to watch for after Jensen Huang’s Japan visit

Jensen Huang left Tokyo with deals spanning Japan's entire tech ecosystem.

Kimi: Threat or menace?
AI News & Artificial Intelligence | TechCrunch

Kimi: Threat or menace?

Chinese company Moonshot AI released a new version of its Kimi model this week, prompting concern about "full AI communism."

Artificial Intelligence - ABC News - Breaking News, Latest News and Videos
"artificial intelligence" - Google News

Artificial Intelligence - ABC News - Breaking News, Latest News and Videos

Artificial Intelligence  ABC News - Breaking News, Latest News and Videos

The latest AI news we announced in June 2026
AI

The latest AI news we announced in June 2026

Here are Google’s latest AI updates from June 2026.

OpenAI is scared of open-weight models. Should the US be?
AI News & Artificial Intelligence | TechCrunch

OpenAI is scared of open-weight models. Should the US be?

Talk of banning Chinese-made open-weight LLMs reveals the challenge of turning AI into a business.

We’re strengthening our presence in Alabama through new investments and community support.
AI

We’re strengthening our presence in Alabama through new investments and community support.

Google has announced a $1.5 billion investment for 2026 and 2027 to expand its data center campus in Jackson County, Alabama. Operating since 2019 on a repurposed former…

AI’s most important protocol is getting a little bit easier to use
AI News & Artificial Intelligence | TechCrunch

AI’s most important protocol is getting a little bit easier to use

Under the new system, the protocol will take a looser, "stateless" approach to session IDs on the server side, similar to how most ordinary websites already work.

Frederick Health Hospital using artificial intelligence to advance medical treatment for patients vulnerable to seizures - DC News Now
"artificial intelligence" - Google News

Frederick Health Hospital using artificial intelligence to advance medical treatment for patients vulnerable to seizures - DC News Now

Frederick Health Hospital using artificial intelligence to advance medical treatment for patients vulnerable to seizures  DC News Now

Neil Rimer thinks the AI money is coming back out
AI News & Artificial Intelligence | TechCrunch

Neil Rimer thinks the AI money is coming back out

Neil Rimer, the venture capitalist who co-founded Index Ventures, predicts the historic wealth AI is generating in Silicon Valley will have to be redistributed, voluntarily or involuntarily.

USC at the Leading Edge of AI - USC Today
"artificial intelligence" - Google News

USC at the Leading Edge of AI - USC Today

USC at the Leading Edge of AI  USC Today

‘Odyssey’ director Christopher Nolan calls AI an obvious ‘Trojan horse’
AI News & Artificial Intelligence | TechCrunch

‘Odyssey’ director Christopher Nolan calls AI an obvious ‘Trojan horse’

"Everybody knows the Greeks are inside."

Safety and alignment in an era of long-horizon models
OpenAI News

Safety and alignment in an era of long-horizon models

OpenAI shares lessons from deploying long-running AI models, highlighting new safety risks, observed failures, and improved safeguards through iterative deployment.

YouTube clarifies policies around AI slop and upsetting videos
AI News & Artificial Intelligence | TechCrunch

YouTube clarifies policies around AI slop and upsetting videos

YouTube has updated its monetization policies to more clearly define the kinds of AI-generated and low-quality videos that can’t earn ad revenue.

Nebraska Snapshot shows AI awareness differs by age, education, income - Nebraska Today
"artificial intelligence" - Google News

Nebraska Snapshot shows AI awareness differs by age, education, income - Nebraska Today

Nebraska Snapshot shows AI awareness differs by age, education, income  Nebraska Today

The latest AI news we announced in May 2026
AI

The latest AI news we announced in May 2026

Here are Google’s latest AI updates from May 2026

How sales teams use ChatGPT Work
OpenAI News

How sales teams use ChatGPT Work

See how sales teams can use ChatGPT Work to create pipeline briefs, meeting prep packets, forecast reviews, account plans, and stalled-deal diagnoses from real work inputs.

London firms warn of widening AI skills gap - BBC
"artificial intelligence" - Google News

London firms warn of widening AI skills gap - BBC

London firms warn of widening AI skills gap  BBC

GPT-Red: Unlocking Self-Improvement for Robustness
OpenAI News

GPT-Red: Unlocking Self-Improvement for Robustness

Explore GPT-Red, OpenAI’s automated red teaming system that uses self-play to improve AI safety, alignment, and prompt injection robustness.

OpenAI Appears to Be Missing Its Sales Goals by a Vast Margin - Futurism
"artificial intelligence" - Google News

OpenAI Appears to Be Missing Its Sales Goals by a Vast Margin - Futurism

OpenAI Appears to Be Missing Its Sales Goals by a Vast Margin  Futurism

5 ways Google Search can level up your thrift and vintage shopping
AI

5 ways Google Search can level up your thrift and vintage shopping

Uncover second-hand scores with AI tools in Google Search and Shopping.

How Google’s A.I. Search Is Imperiling the Open Web - The New York Times
"artificial intelligence" - Google News

How Google’s A.I. Search Is Imperiling the Open Web - The New York Times

How Google’s A.I. Search Is Imperiling the Open Web  The New York Times

5 ways to learn with study notebooks in the Gemini app
Gemini

5 ways to learn with study notebooks in the Gemini app

Study notebooks is a new space in the Gemini app that serves as an interactive learning tool tailored to any student's goals.

The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway
AI | VentureBeat

The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway

Across 157 enterprises, organizations are granting AI agents more autonomy while trusting the evaluations meant to gate that autonomy less. Half have already shipped an agent that passed their internal evaluations and then failed a customer in production; only one in twenty fully trusts automated evaluation today; and the most-cited weakness is that evaluations do not align with real-world outcomes. Yet two-thirds already allow, or are actively engineering toward, deploying agent changes to production on automated evaluation alone — with no human in the loop. The result is an evaluation gap — the distance between how much autonomy enterprises are handing their agents and how far they trust the tests that are supposed to catch the failures. This wave of VentureBeat Pulse Research examines how technical leaders measure agent performance: which reliability and evaluation platforms they use, how they select and trust them, what breaks in production, and how far they are willing to let agents run without a human in the loop. The central finding is an evaluation gap — the distance between the autonomy enterprises are granting their agents and the trust they place in the evaluations meant to govern it. Half of organizations (50%) have, in the past year, deployed an agent or LLM feature that passed their internal evaluations and then caused a customer-facing failure, and a quarter have seen it happen more than once. Trust in the tests themselves is thin: only 5% say they fully trust automated evaluation today, and the single most-cited limitation is that evaluations align poorly with real-world outcomes (29%). Enterprises are discovering that a passing eval is not the same as a working agent. What makes the gap consequential is the direction of travel. Two-thirds of organizations (66%) already permit fully automated, zero-human-in-the-loop deployment for low-risk agents (34%) or are actively engineering their pipelines to allow it within twelve months (33%). At the same time, the evaluation stack that would have to earn that trust is fragmented and immature: the most common primary tools are the model providers’ native evals, tied with having no dedicated tooling at all (17% each); and only about a quarter of enterprises run real-time quality checks on live production traffic. The autonomy is arriving faster than the assurance. Methodology VentureBeat fielded this survey as part of its ongoing Pulse Research series, this survey — the Agentic Reliability & Evals tracker — focused on how technical leaders evaluate agent performance and reliability. Responses are filtered to organizations with 100 or more employees (n=157), drawn from a single survey in June 2026; because this is one wave rather than a pooled multi-month sample, the report reads cross-sectionally and does not infer month-over-month trends. Where questions were multiple-select, those shares can sum to more than 100%. By role the sample is senior and buyer-credible: 38% are final decision-makers for AI purchases and another 34% recommenders or influencers. Product and program managers (15%), consultants and advisors (10%), directors of engineering/IT (8%), and CIOs/CTOs/CISOs (8%) lead the named titles, alongside a large “Other” function (37%). By organization size the sample is mid-market-weighted: 100–499 (37%) and 500–2,499 (27%) employees lead, with 2,500–9,999 (20%), 10,000–49,999 (10%), and 50,000+ (6%) above them. Technology/Software is the largest industry at 23%, followed by Retail/Consumer (15%), Healthcare/Life Sciences (12%), and Manufacturing (10%). At 157 respondents the sample is large enough to read directionally but should be treated as a directional signal rather than a precise measurement; it is self-selected and is not a probability sample. It skews toward the mid-market, so it is best read as the view from organizations actively standing up agent evaluation practices rather than from the largest operators. Note: This survey was rebuilt for the June wave from the earlier “LLM observability and evaluations” survey; because the questions and sample differ, no comparisons are made to the April–May data. Finding 1: A passing eval is not a working agent Half have shipped an agent that passed evals, then failed a customer We asked whether, in the past 12 months, organizations had deployed an agent or LLM feature that passed their internal evaluations but then caused a customer-facing failure. Half of those that run evaluations had. This is the report’s defining number. Half of organizations (50%) have shipped an AI feature that cleared their internal evaluations and then failed in front of a customer — an incorrect output, a broken workflow, or a quality incident — and a quarter have seen it happen more than once. Only 36% report no such failure, and the remainder either run no pre-deployment evaluations (8%) or don’t track the root cause closely enough to know (6%). The failure is precise and expensive: the evaluation said the agent was ready, and it was not. Everything that follows — how enterprises trust their evals, what they monitor, and how much autonomy they grant — is shaped by this experience. Finding 2: Almost no one fully trusts automated evaluation The top complaint: Evals don't match real-world outcomes We asked which limitation most reduces trust in automated agent evaluations today. Only a sliver of enterprises had no complaint at all. Trust in automated evaluation is scarce, and specific. Only 5% of organizations say they fully trust automated evaluation as it stands — meaning 95% name a limitation that holds them back. The most common, at 29%, is the one that most directly explains Finding 1: evaluations align poorly with real-world outcomes, passing agents that later fail. Bias or inconsistency (21%) and a lack of explainability (18%) follow — enterprises cannot always tell why an evaluation reached its verdict — and 17% cite data-leakage or privacy concerns in the evaluation process itself. The tests meant to certify agents are not yet trusted to certify them, which is precisely why the autonomy trajectory in Finding 3 is so striking. Finding 3: The autonomy ceiling is rising anyway Two-thirds already allow, or are building toward, zero-human deployment We asked whether organizations would let an autonomous agent deploy a code or system change to production on automated evaluation results alone, with no human-in-the-loop validation. The trajectory runs straight through the trust gap. Here is the paradox at the heart of the report. Even though almost no one fully trusts automated evaluation (Finding 2), two-thirds of organizations (66%) either already allow zero-human-in-the-loop deployment for low-risk agents (34%) or are actively engineering their pipelines to permit it within a year (33%). Only 22% rule it out for the foreseeable future. The direction is unambiguous: enterprises are moving to let evaluations gate production autonomously — removing the human check — at the same moment they say those evaluations don’t reliably match reality. The autonomy ceiling is rising faster than the assurance beneath it, which is the mechanism by which the false-confidence failures of Finding 1 will scale rather than shrink. Notably, the autonomy bet is not just a small company phenomenon. Splitting the sample by company size, larger enterprises are slightly further down the path toward zero human review than smaller companies (70% versus 64%) and slightly more likely to have shipped an evaluation-passing agent that then failed a customer (54% versus 48%). The assumption that large, regulated organizations are holding the human in the loop longest is, in this sample, backwards.  To be sure, these are directional figures, since the survey was not a huge sample — 57 respondents from companies with 2,500+ employees and 100 from companies smaller than that.  Finding 4: The evaluation stack is fragmented and provider-led Provider-native evals lead — tied with no dedicated tool at all We asked which agent reliability or evaluation platform enterprises primarily use today. The market has no clear leader — and a large share has nothing dedicated. The evaluation layer is early and unconsolidated. Provider-native tooling leads — OpenAI’s native evals and traces (17%) and Anthropic’s Claude Console evals (13%) together outweigh any independent platform — but it is tied at the top by a striking answer: 17% of enterprises use no dedicated agent-evaluation tooling at all, a notable gap for organizations shipping agents to customers. The specialist evaluation vendors — DeepEval (12%), Braintrust (8%), LangSmith, Weave, Promptfoo, Langfuse, Arize — are scattered across single to low double digits, and 11% have built their own. No independent platform has yet become the category standard, which leaves most enterprises evaluating agents with provider-native tools, home-grown scripts, or nothing. Finding 5: Production monitoring rarely watches output quality Only a quarter run real-time quality checks on live traffic Production monitoring for an AI agent can watch two very different things. It can watch whether the system is functioning — is the agent up and responding, did each request complete, how fast, at what cost, with any errors. Or it can watch whether the agent's output is correct — automated checks that evaluate the content of each answer as it goes out: did the agent give the right answer, take the right action, stay within policy. The distinction matters because a confidently wrong answer is invisible to the first kind of monitoring: the request completes, the response is fast, no error is thrown, and every functioning-metric reads healthy. We asked organizations which kind their live production monitoring is built for today. Grouped by what is actually being watched, the split is stark: 51% of organizations monitor only whether the agent is functioning, while 23% monitor whether its answers are right. Counting the ad-hoc reviewers and the don't-knows, roughly three-quarters of organizations run no automated, real-time evaluation of output correctness in production — they can see that the system is up and what it costs, and they are taking the correctness of its answers on faith. That blind spot is the runtime counterpart to the pre-deployment gap in Finding 1: the same organizations engineering the human out of the deployment decision mostly cannot see, in real time, when the deployed agent starts getting things wrong. Finding 6: Bought on cost, measured on consistency Price and integration drive selection; evaluation consistency is the goal We asked what most influenced enterprises’ choice of an evaluation vendor, and what they treat as their primary measure of success. Both answers are pragmatic. Enterprises buy evaluation tooling on economics and trust it on repeatability. Cost of evaluations (28%) narrowly leads selection, just ahead of ease of integration (27%) and evaluation accuracy (24%) — breadth of observability (13%) and vendor roadmap (4%) matter far less. On what success looks like, more than a third (36%) name evaluation consistency — getting the same verdict on the same behavior every time — well ahead of speed of experimentation (19%), reduction in failures (18%), production visibility (13%), and compliance (11%). The emphasis on consistency is telling: before enterprises can trust an evaluation’s verdict, they need it to be stable — the very property whose absence (bias and inconsistency) ranked among the top trust limitations in Finding 2. Satisfaction with current tooling is only moderate, averaging 3.8 on a five-point scale across overall satisfaction, ease of implementation, and value for money. Finding 7: The next dollar goes to humans and observability Investment is flowing to oversight, not just automation We asked which reliability and evaluation investment will grow most over the next year. The money is going toward watching agents more closely — including with people. The second-largest planned investment — behind only production observability — is human review workflows, at 26%. Read against Finding 1, that is the report's quietest contradiction: at the same moment two-thirds of enterprises are engineering the human out of the deployment decision, more of them plan to grow spending on human reviewers (26%) than on the automated evaluation pipelines (16%) that would replace them. The zero-human trajectory and the human-review budget are rising in the same companies at the same time. Indeed, only 8% report that their budget is not increasing. Taken together, enterprises are hedging: building toward autonomy while spending to watch agents more closely and keep humans available for the calls that automated evaluation cannot yet be trusted to make. Finding 8: A tooling reshuffle is coming Nearly two-thirds plan to adopt or switch platforms within a year We asked whether enterprises plan to adopt a new, additional, or replacement evaluation platform, and which they are considering. Few intend to stand pat. The evaluation market is wide open. While 36% have no plans to change, a clear majority (64%) intend to adopt a new, additional, or replacement platform within twelve months, and 31% within the next quarter. The consideration set points where current usage is thinnest: Confident AI’s DeepEval leads what enterprises are evaluating (20%), ahead of OpenAI’s native evals (13%) and Braintrust (9%) — the open-source specialists drawing more interest than their present footprint. Given that so many enterprises today rely on provider-native tools or nothing at all (Finding 4), this is less a defection than a first real wave of tooling adoption — the moment the evaluation layer starts to consolidate. Which platforms earn that trust, in a market where almost no one trusts automated evaluation yet, is the open question this series will keep tracking. The bottom line: An evaluation gap that autonomy will widen, not close Organizations with 100 or more employees are granting AI agents more independence than they trust their evaluations to support. Half have already shipped an agent that passed its evals and then failed a customer; almost none fully trust automated evaluation, chiefly because it doesn’t match real-world outcomes; and most watch production for uptime and cost rather than for whether the agent’s answers are right. Yet two-thirds already allow, or are actively building toward, deploying to production on automated evaluation alone. The vendor market is early and unsettled: the most common primary evaluation tools are provider-native evals, tied with no dedicated tooling at all, and a clear majority plan to adopt or switch platforms within the year. Encouragingly, the next dollar is going to observability and — pointedly — human review, suggesting enterprises sense the gap even as they engineer past it. At 157 respondents in a single wave this is a directional read, skewed toward the mid-market — but the direction is clear: autonomy is being granted on the strength of evaluations that the people granting it do not yet trust. The evaluation gap is not a coverage problem that more tests alone will close; it is a problem of evaluations that reflect reality and can be trusted to gate it. The open question for later waves is whether assurance catches up to autonomy — or whether the false-confidence failures move from customer incidents into changes that deploy themselves. Based on survey responses from 157 qualified enterprise respondents (100+ employees), drawn from a single June 2026 wave. This is a directional read rather than a precise measurement — the sample is self-selected, not a probability sample, and skews toward the mid-market. Respondents include product and program managers, consultants and advisors, directors of engineering/IT, and CIOs/CTOs/CISOs, among other functions, across technology/software, retail/consumer, healthcare/life sciences, manufacturing, and other industries.

Unlocking Britain’s next era of productivity: Building a nation of AI trailblazers
AI

Unlocking Britain’s next era of productivity: Building a nation of AI trailblazers

Google UK shares its latest Economic Impact Report and how to enable more people to unlock the benefits of AI-powered technologies.

Getting started with ChatGPT
OpenAI News

Getting started with ChatGPT

Learn how to use ChatGPT, start your first conversation, and discover simple ways to write, brainstorm, and solve problems with AI.

Can an Apple lawsuit derail OpenAI’s hardware plans?
AI News & Artificial Intelligence | TechCrunch

Can an Apple lawsuit derail OpenAI’s hardware plans?

On the latest episode of Equity, we debate whether Apple's lawsuit will cast a shadow over OpenAi's much-discussed plans to get into hardware and go public.

AI is writing, acting and producing China’s minidramas. It’s shaking a $14 billion industry. - NBC News
"artificial intelligence" - Google News

AI is writing, acting and producing China’s minidramas. It’s shaking a $14 billion industry. - NBC News

AI is writing, acting and producing China’s minidramas. It’s shaking a $14 billion industry.  NBC News

The 6 best vibe coding tools in 2026
The Zapier Blog

The 6 best vibe coding tools in 2026

Decades ago, building the Facebooks of the world was reserved for a small elite clad in technical skills. Today, with a sequence of good prompts, endless curiosity, and good testing practices, you too can stand next to the big names and launch your own app—and no, you don't need to know how to write a function. That's what vibe coding is all about: as coined by Andrej Karpathy, you build an app using natural language, and forget that the code is even there. While the hype is mostly gone—the firs

Head of US AI safety agency resigns - Reuters
"artificial intelligence" - Google News

Head of US AI safety agency resigns - Reuters

Head of US AI safety agency resigns  Reuters