Beyond the Big Three: Challengers, Safety Gaps, and the Expanding Reach of AI Agents
August 23, 2026 • 10:42
Audio Player
Episode Theme
Beyond the Big Three: Challengers, Safety Gaps, and the Expanding Reach of AI Agents
Sources
Transcript
Alex:
Good morning, everyone, and welcome back to Daily AI Digest! It's August 23, 2026, and I'm Alex.
Jordan:
And I'm Jordan. We've got a jam-packed show today — a scrappy underdog AI lab claiming to beat the giants, some uncomfortable questions about what happens if a model goes rogue, OpenAI doing a total policy about-face, AI avatars grading your business pitch, and yes, we're talking about the new Pixel phone.
Alex:
It's a lot. But before we dive in, I have to mention this hibernating mice story I saw — apparently when mice hibernate, they lose a bunch of synapses but somehow still remember stuff.
Jordan:
Honestly, that's basically what happens to me every Monday morning before coffee, and I still remember your birthday, so.
Alex:
Low bar, but I appreciate it. Okay, speaking of things retaining knowledge they maybe shouldn't be able to — let's talk about this AI lab claiming to out-research the research giants.
Jordan:
Right, this is a fun one. According to TechCrunch, there's a British startup called Inherent, founded by former DeepMind folks, and they're saying their AI agent — it's called Faraday — actually outperformed both Anthropic and OpenAI models at replicating scientific research papers.
Alex:
Wait, replicating research papers? Like, redoing someone else's experiment to see if you get the same result?
Jordan:
Exactly. It's a huge deal in science because replication is expensive, slow, and honestly kind of thankless — nobody gets a Nobel Prize for confirming someone else's finding. So if an AI agent can reliably replicate studies, that's a massive unlock for speeding up how we validate scientific claims.
Alex:
And this is a company nobody's really heard of beating OpenAI and Anthropic? That feels like a big deal, or is it more of a 'well, it depends how you define winning' situation?
Jordan:
That's the right instinct to have. Benchmarks in AI are notoriously squishy — who picked the papers, what counts as a successful replication, was the test set public beforehand. So yes, it's exciting, but it's also very possible Inherent picked a benchmark that plays to Faraday's strengths.
Alex:
So it's less 'David slays Goliath' and more 'David picked a fight David was pretty confident about.'
Jordan:
Ha, that's fair. But even with that caveat, it's a signal worth watching. It shows agents are moving beyond writing code or answering chat questions into actual scientific workflows — reading a paper, understanding the methodology, running the analysis, checking if the numbers hold up.
Alex:
That's genuinely useful for R&D teams, though, right? Not just a cool demo.
Jordan:
Very much so. If you're a pharma company or a university lab, having an agent that can pressure-test published findings before you sink six months of budget into building on them — that's real value, assuming the claims hold up under scrutiny.
Alex:
Okay, so cautiously optimistic, pending someone else running the same test.
Jordan:
That's exactly where I'd land. Which, funny enough, ties into our next story, because it's all about what happens when we can't fully verify or control what these AI systems are doing.
Alex:
Ooh, ominous transition. What have we got?
Jordan:
So this one's also from TechCrunch, and it's a bit unsettling. A new study found that the leading frontier AI labs have basically no publicly documented plan for how they'd contain a rogue or misbehaving AI model.
Alex:
Hold on — none? Like, zero written-down 'break glass in case of emergency' plan?
Jordan:
Essentially, yes, at least publicly. The study looked at how prepared these labs claim to be versus what they've actually documented, and there's a pretty big gap between 'we take safety seriously' marketing language and any concrete, specific containment protocol.
Alex:
What would 'containment' even look like in practice? Like unplugging a server?
Jordan:
It's more nuanced than that — think kill switches, sandboxing environments so an agent can't access the wider internet or take real-world actions, monitoring for specific warning behaviors, and having a clear chain of command for who pulls the plug and when. The concern is that as models get more capable and more agentic — remember, these things are now writing code, executing tasks, even doing scientific research like we just discussed — the potential blast radius of something going wrong keeps growing, but the safety planning hasn't kept pace.
Alex:
Did the study call out specific labs, or is everyone equally in the dark?
Jordan:
There's variation — some labs scored better than others in terms of having at least some documentation — but the overarching finding is that the industry as a whole is under-prepared relative to how fast capabilities are advancing. Nobody's acing this test.
Alex:
That's genuinely a little scary, especially with agents like Faraday running around doing autonomous research tasks.
Jordan:
Right, and it matters a lot for anyone building products on top of these models. If you're a company integrating GPT or Claude or Gemini into your product, you're inheriting some of that risk. Knowing whether your provider has a real incident response plan should be part of your due diligence, the same way you'd check a cloud provider's security certifications.
Alex:
So this isn't just an academic worry, it's a genuine enterprise risk question.
Jordan:
Exactly, and it dovetails nicely into our next story, because it's about the regulatory side of this exact problem.
Alex:
Let me guess — someone's finally trying to write some actual rules?
Jordan:
Sort of, and the twist is who's asking for it. According to TechCrunch, OpenAI is now urging California to strengthen SB 53, which is an AI safety bill that OpenAI previously opposed.
Alex:
Wait, they opposed it and now they want it to be stronger? That's a full one-eighty.
Jordan:
It really is. SB 53 would add things like mandatory safety testing disclosures and incident reporting requirements for frontier model developers. OpenAI's initial position was pretty standard industry pushback — too burdensome, could slow innovation, that sort of thing.
Alex:
So what changed? Did they suddenly have a safety epiphany?
Jordan:
That's the big question, and honestly the article raises it too — is this genuine safety concern, competitive positioning, or just good PR timing given everything we just talked about with the containment study?
Alex:
It does seem convenient that this comes right after a report saying labs have no rogue-AI containment plan.
Jordan:
Right, the timing is notable. There's also a competitive angle — if OpenAI supports stronger regulation now, it could set rules that are harder for smaller, less-resourced startups, like our friends at Inherent from story one, to comply with. Established players sometimes like regulation because they can afford compliance costs that smaller competitors can't.
Alex:
Oh, that's a cynical take, I like it. So it could be less 'we've seen the light' and more 'let's build a moat out of paperwork.'
Jordan:
It's possible, though it's probably not that simple either. Anthropic has generally been more openly supportive of safety regulation from the start, and Google has been more cautious and measured in its public stance. So this shift brings OpenAI's public posture more in line with Anthropic's, which is interesting in itself.
Alex:
What would this actually mean for developers using these models day to day?
Jordan:
If SB 53 passes with teeth, you could see more transparency around model capabilities and safety testing before releases, possibly slower release cadences as labs comply with reporting requirements, and clearer incident disclosure if something does go wrong in production. For developers building on these platforms, that's honestly a net positive for predictability, even if it means fewer surprise model drops.
Alex:
Interesting. Okay, from lawmaking to something a little more, let's say, personal — Harvard's AI avatars?
Jordan:
Yes! This one's fun. According to TechCrunch, Harvard Business School runs this $699 startup bootcamp called HBS Foundry, and they're now using AI avatars of their actual instructors to give feedback during mock pitches and board meetings.
Alex:
Wait, so it's not just a generic AI coach, it's like a digital clone of a specific real professor?
Jordan:
Exactly, they've built avatars modeled on real instructors, so when you're practicing your startup pitch, you get feedback that's styled after that person's actual expertise and mannerisms, presumably trained on their past lectures, writing, feedback patterns.
Alex:
That's wild. Is it actually good feedback, or is it more like a fancy chatbot wearing a professor costume?
Jordan:
That's the open question, honestly. The value proposition is scale — one instructor's expertise, available to give feedback to hundreds of students simultaneously, at any hour, without them getting tired or repeating themselves. But there's a real authenticity question. Are you getting genuine expert insight, or a plausible-sounding simulation of it?
Alex:
I feel like for something as high-stakes as a board meeting simulation, you'd want the real nuance a human catches — like if you're bombing the room and don't realize it.
Jordan:
Right, and that's the tension. Simulated pressure can be genuinely useful for building reps and confidence, low stakes practice before the real thing. But subtle read-the-room stuff, or catching when a founder's story doesn't quite add up emotionally — that might still need a human in the loop.
Alex:
It's a bit like flight simulators, though, right? You wouldn't want your first flight to be a passenger plane, but simulator hours still make you a better pilot.
Jordan:
That's a great analogy, actually. And this fits a bigger pattern we're seeing — AI agents moving into coaching and mentorship roles well beyond coding assistants. We've talked about AI tutors, AI therapists in past episodes, and now AI business mentors. The common thread is using foundation models to make expert-level guidance scalable, for better or worse.
Alex:
Would you pay $699 to pitch your startup idea to a professor-bot?
Jordan:
Honestly, if it helps me not freeze up in front of real investors later, sure, sign me up. Speaking of paying for things that might be more marketing than substance, let's talk about the new Pixel.
Alex:
Oh, I saw this — the Pixel 11 Pro XL review. Not exactly glowing?
Jordan:
According to TechCrunch, the reviewers found the cameras snappier, the AI features present, but overall it's a fairly iterative upgrade rather than a groundbreaking one. There's a new AI feature called Rambler.
Alex:
Rambler — okay, that name is doing a lot of work. What does it actually do?
Jordan:
The details are a bit thin in the coverage, but it seems to be some kind of conversational or assistant feature layered into the phone's core experience — think along the lines of proactive suggestions or a more natural voice interaction, building on what Gemini's already doing on-device.
Alex:
So is it genuinely useful, or is it dressing up a modest hardware refresh?
Jordan:
That's the real question the reviewers are wrestling with. The pattern we're seeing across the industry right now is that when the hardware gains are marginal — slightly better sensor, slightly better chip — companies lean hard on 'but look at the AI' to justify the upgrade cycle.
Alex:
It's basically the software equivalent of 'now with 10% more sparkle.'
Jordan:
Ha, exactly. And there's a real risk of consumer fatigue here. If every single flagship phone announcement leads with some new AI feature name, at some point people stop caring, or worse, get suspicious that it's just marketing filler.
Alex:
Do you think that fatigue is already setting in?
Jordan:
I think we're getting close. The features that stick are the ones that solve an actual annoying problem, better call screening, genuinely useful photo editing, real-time translation that works. The ones that feel bolted on for the keynote slide tend to fade from memory fast.
Alex:
So Rambler's fate depends on whether people are still talking about it in six months.
Jordan:
Pretty much. Ask me again at Pixel 12 season.
Alex:
Fair enough. Okay, that is a lot of ground covered today — from a scrappy UK lab challenging the AI giants, to labs not having a rogue-AI game plan, to OpenAI's regulatory flip, AI professors, and AI-flavored phone cameras.
Jordan:
It really shows how the AI story right now isn't just about the big three labs anymore — it's challengers nipping at their heels, safety questions nobody's fully answered, and AI agents quietly showing up in classrooms, phones, and research labs everywhere.
Alex:
Beyond the big three indeed. Thanks for hanging out with us today, everyone.
Jordan:
We'll be back tomorrow with more AI news. Until then, this has been Daily AI Digest — take care, everybody.