1009 | AI Agents Everywhere: Desktop, Terminal, Cloud and Beyond

||Download

Show notes

A fast tour of this week's launches: AI agents that act on your behalf across desktop, terminal, cloud, and phone, plus the models and infrastructure powering them, and a few non-AI gems.

Timeline

  • 00:00:04 Opening
  • 00:00:48 Managing and orchestrating AI agents
  • 00:06:07 AI agents that do real work: terminal, cloud and office
  • 00:10:17 New models: small, cheap, and scarily human
  • 00:13:08 Infrastructure, routing and self-hosted alternatives
  • 00:17:24 Personal, creative and privacy tools
  • 00:21:14 Judgment and human oversight
  • 00:23:11 Hardware, safety and fun
  • 00:26:21 Closing

Related links

This episode is produced by Bri. Bri uses advanced AI technology to turn the feeds you care about into podcasts made for listening. Contact us at hi@bri.so.

Transcript

Mia: Welcome back to the show, everyone. I'm Mia.

Milo: And I'm Milo. And honestly, if you had to pick one thread running through everything this week, it's that AI agents have stopped being demos and started being... managed. Like, actual infrastructure with supervisors, approvals, control panels.

Mia: Right — and we're going to work through that, from the ways people are herding these agents, to the models behind them, the infrastructure underneath, the personal and creative stuff, and yes, a pair of motorized shoes at the end because not everything is an agent. Let's start with the orchestration story, because it's the most interesting shift.

Milo: So picture the typical setup right now. You've got coding agents like Codex or Claude Code running on your machine, and they eat your screen. You can't work while they work. That's the problem Offroad attacks, and it's a clever angle. Offroad gives your coding agent a second macOS account that's already logged in, dedicated to GUI work. So the agent can click around in actual applications while your own screen stays free. It's MIT licensed and it's Apple Silicon only.

Mia: I like that it doesn't pretend the GUI doesn't exist. A lot of agent tooling is terminal-only, and then someone eventually has to open a browser or an app and the whole flow breaks. Offroad is basically saying: fine, give the agent its own desktop so it never touches yours. That's a really concrete user problem — "I want to delegate, but I still need my computer."

Milo: And there's a real open question there, which is how well agents actually handle GUI work in practice. The maker is claiming the setup works; whether agents are reliable enough clicking through apps day to day, that's unproven territory. But the mechanism itself is clean.

Mia: Now, once you have agents running — maybe several — you need to manage them. And that's where BotBus comes in. BotBus lets you manage your coding agents from your phone or your Apple Watch. So Codex, Claude Code, whatever you're running — you can check on them, and crucially, approve things, from away from your desk.

Milo: That approval part is the key detail. These agents take actions, and somebody has to sign off. BotBus does that, and it's end-to-end encrypted, which matters because you're essentially sending control of a machine that has your code over the network. It supports macOS, Linux, and Windows, so it's not locked to one platform.

Mia: So Offroad solves "the agent is hogging my computer," BotBus solves "I need to supervise the agent when I'm not at the computer." Different problems, complementary products. And then there's a third layer: what happens when these sessions run for a long time?

Milo: That's pmtui. It's a supervisor written in Rust, built on tmux, for long-running Claude and Codex sessions. The interesting idea is the autopilot: routine decisions get resolved automatically, and only the things that genuinely need a human get escalated to you. So instead of babysitting every little step, you're only pulled in when it matters.

Mia: That's a design philosophy showing up everywhere this week, by the way — AI proposes, humans approve the important stuff. Keep that in your head, because it comes back later.

Milo: One thing worth flagging on pmtui: "routine" is doing a lot of work in that sentence. What counts as routine? How does the autopilot decide what's safe to resolve on its own? Those thresholds are where this either works or causes headaches. There's no evidence in what we have about how it calibrates that.

Mia: Okay, so we've got Mac-specific stuff and a cross-platform remote manager. What about Windows users?

Milo: Clippo. It's a visual canvas on Windows for orchestrating agents — Claude Code, Codex, and others. So instead of terminal windows, you lay out the work visually. And two details stand out: memory is stored locally, and there's no cloud and no login. It's free on the Microsoft Store.

Mia: The no-cloud, no-login thing is genuinely notable, because a lot of agent orchestration tools want to be a service. Clippo keeping everything local lowers the trust barrier a lot — your agent's context and memory never leave your machine. That's a meaningful differentiator if you're wary of sending project context to yet another server.

Milo: And then the most ambitious one in this group: OpenSwarm. It's a desktop app for Mac where swarms of agents work in parallel with shared context — they've got a browser and apps to work with. It's free, but there's a waitlist.

Mia: That shared-context point is the differentiator. Running five agents in parallel isn't hard; having them actually share what they know while working is the hard part, and that's what OpenSwarm is claiming. But here's the honest caveat: it's waitlisted. Pricing beyond "free for now" is unknown, and we don't have adoption details or evidence of how well the swarm coordination actually performs. Treat it as a promising claim, not a proven result.

Milo: Right. So the picture from this whole group: agents are getting a management stack. Second accounts, remote approvals, supervisors with escalation logic, visual canvases, parallel swarms. That's what infrastructure for a new kind of worker looks like when it's being built in public.

Mia: And once you accept agents are infrastructure, the natural question is: what are these agents actually doing all day? Which brings us to agents that do real work — in the terminal, in the cloud, and in your office suite.

Milo: Start with Nova CLI. It's an AI developer that lives in your terminal. It reads your project, edits code, runs tests, runs builds. The full loop. And the pricing model is interesting: paid plans include access to 160+ models, and you don't need to bring your own API keys.

Mia: That last bit is a bigger deal than it sounds. Managing API keys across providers is a real hassle and a real cost-surprise risk. If Nova bundles the models into the plan, they're competing on convenience. The trade-off, of course, is you're trusting their plumbing and their pricing rather than paying providers directly. For a developer who just wants the thing to work, that's probably a good trade.

Milo: Then there's Hark Pro, which is aiming at a much broader target — a personal AI that acts proactively. Email, calendar, shopping, travel. And the architecture is the striking part: it comes with your own computer in the cloud, and your credentials are encrypted. Plus actions are visible — you can see what it's doing.

Mia: "Visible action" is the trust mechanism there. An AI that books your flights invisibly is terrifying; one where you can watch what it did is at least auditable. But let's be careful: proactive action on your email and your wallet is the highest-stakes version of this, and the claims here are the maker's claims. Whether it reliably judges "should I buy this" or "should I reply this way" — that's exactly the kind of thing that needs real-world evidence before you hand over credentials.

Milo: Moving into the office: Claude for Google Workspace. This is Anthropic putting Claude literally inside Docs, Sheets, and Slides. It's in beta, on paid plans. It can edit files from within those apps, and you can also create and edit files starting from Claude itself.

Mia: That's the "agent in the document" pattern. Versus copy-pasting between a chat window and a doc, having the model operate in place removes a whole class of friction. The beta label and the paid-plan gate are the practical limits — availability isn't universal yet.

Milo: Then there's a couple of more specialized ones. Polylane for Vercel is an on-call AI engineer through the Vercel Marketplace. When there's a regression, it investigates — with evidence — and opens a pull request with the fix for a human to review.

Mia: Note the shape of that: it doesn't merge, it opens a PR for review. Again, propose, human approves. That's consistent with everything we've said. And it's scoped tightly — regressions — which is smart. A narrow, well-defined job is where this kind of automation actually works.

Milo: And Leanback is aimed at engineering managers, living in Slack. Daily check-ins, it reads tickets and calendars, and produces action plans. There's a 30-day trial.

Mia: Engineering managers are an underserved group for tooling — they're drowning in status aggregation. If Leanback genuinely synthesizes tickets plus calendars into something useful, that's real time back. Whether its read of the situation is accurate is the open question, as always.

Milo: One more piece of context for this section: Kloudmate 2.0, which we'll come back to later, fits here too — it's an observability platform that's added AI agents for root-cause analysis and dashboards. So even monitoring is getting agents now.

Mia: Okay. So we've got agents doing terminal work, office work, ops work. All of that runs on models. And the model story this week is: small, cheap, and unsettlingly human.

Milo: Claude Haiku 5.5 first. This is Anthropic's small model tier, and the pitch is faster and about 75% cheaper. Notably, it's the first Haiku with an effort adjustment — you can dial how hard the model works. It's aimed at high-volume use.

Mia: Effort adjustment is the underrated feature there. If you're running thousands of agent calls — which, given everything we just discussed, people are — being able to say "this call is trivial, use minimum effort" is where that 75% cost reduction actually materializes. Small cheap models are what make the agent-everywhere vision economically possible. That's the connection.

Milo: Then the creepy end of the spectrum: Griffin, from Tavus. It's described as the first duplex video-to-video human interaction model — meaning real-time, back-and-forth, face-to-face over video. And in testing, 48% of participants thought they were talking to a real person.

Mia: Sit with that number for a second. Nearly half of people on a live video call couldn't tell. That's a genuine milestone and a genuine concern at the same time. The plausible-use story is things like avatars, customer-facing video agents. But it also means the "is this a person on the call" question just got much harder. The fact that it's in preview for testers — not broadly deployed — is worth noting. We don't know how it behaves at scale or what guardrails exist in the wild.

Milo: And then the evaluation side: Cekura Bench. This is a speech-to-speech benchmark that ran nine models through real phone calls across 82 scenarios. Among single models, GPT Realtime 2.1 scored best at 79.3 percent. But here's the twist — a cascade, meaning chaining models, beat every single one at 82.9 percent.

Mia: The cascade result is the interesting finding. It suggests that for voice, the frontier isn't one model getting better — it's routing between models, using the right one for each moment. Which is a perfect segue, actually, because routing is exactly what the infrastructure layer is doing.

Milo: Right — and that's our next stop. Infrastructure, routing, and self-hosted alternatives.

Mia: Liquid Inference is an LLM router with a specific economic idea: providers compete on price for every prompt. So each request, there's effectively an auction, and you get the cheaper qualifying option. The APIs are compatible with OpenAI and Anthropic formats, so it's a drop-in. And there's $20 in free credits for the first 500 users.

Milo: Connect that to the Cekura finding — if cascades and routing beat single models anyway, and if providers are competing per-prompt on price, then the smart architecture is: never hardcode one model. The router is becoming a layer of the stack. The open question is how "compete on price" interacts with quality guarantees — cheapest isn't automatically good enough for your workload, and we don't have details on how that matching works.

Mia: Now the self-hosted side. Tide is a Go-based alternative to Jitsi and Zoom, where the media stays on your own server. It has server-side recording and OIDC login, and it's open source.

Milo: Media on your own server is the whole point — for a lot of organizations, "our calls never touch someone else's cloud" is a requirement, not a preference. Recording server-side also solves the quality problem that plagues client-side recording. The catch with all self-hosted stuff is the operational burden — you're the ops team now. But that's the trade you choose.

Mia: Tunnl.gg addresses a smaller but very common pain: exposing your localhost to the world. It uses SSH, gives you a stable URL — your SSH key is your identity — requires no installation, and gives you HTTPS and live logs. The server is open source, MIT.

Milo: Stable URL plus key-as-identity is the nice part. Traditional tunneling means new URLs each time or account management. Here your existing SSH key is the whole auth story. Good fit for demos, webhooks, that kind of thing.

Mia: Then two agent-ready infrastructure plays. OpenSEO is an open-source alternative to Semrush, with MCP support — meaning AI agents can use it directly — with data from DataForSEO. It's sitting at 22.7 thousand stars on GitHub.

Milo: Careful with that star count — it tells you people are interested, not that the data quality rivals Semrush. SEO data is only as good as its source, and Semrush's moat has always been data. OpenSEO's real differentiator is the MCP angle: SEO work is exactly the kind of thing people want agents to do, and being agent-callable from day one is the positioning. Whether the underlying data holds up is the open question.

Mia: Kloudmate 2.0, which we mentioned earlier: full-stack observability, native OpenTelemetry, now with AI agents — Assistant, Builder, Investigator, Docs — aimed at root-cause analysis and dashboard building.

Milo: The named roles are interesting. Instead of one generic chatbot, you get an Investigator for RCA and a Builder for dashboards. OpenTelemetry-native matters too, because it means it ingests the standard instrumentation you already have.

Mia: And rounding out the infrastructure group, something small and domestic: Udon, a control panel for Mac home servers on Apple Silicon. Containers, storage, terminal, and an AI assistant. $29 lifetime, with a five-day trial.

Milo: Twenty-nine dollars lifetime is a friendly price for home-server folks — that's a hobbyist audience, and lifetime pricing fits how they think. The AI assistant on a home server panel is the little glimpse of where this is all heading: agents managing even your homelab.

Mia: Okay, we've done agents, models, infrastructure. Time to come back down to earth: personal, creative, and privacy tools.

Milo: Let's start with the privacy one, because it's the most pointed. Off the Record is a macOS app that changes your voice in real time — specifically to block AI transcription by things like meeting notetakers and copilots. And the processing is 100 percent local.

Mia: This is a direct counterpunch to the wave of AI meeting recorders. The premise: if these tools transcribe everything by default, individuals deserve a way to opt out. Real-time voice alteration that defeats the transcription, entirely on-device, so your audio never even goes to a third party. The limitation is obvious, though — if you speak in a meeting, altered, you're also changing how humans hear you. There's a real tension between privacy and just... being understood.

Mia: How well it works against specific transcription tools is a claim we can't verify from here.

Milo: Next, Revela, which is an iPhone app that's almost anti-AI in spirit. It's a film camera: twelve films to choose from, no preview, and you wait — an hour to days — before your photos are revealed. $7.99 one-time, no account, works offline.

Mia: The wait is the product. It's recreating the emotional mechanics of film photography — the delay, the uncertainty — in software. One-time pricing and no account is refreshingly respectful of the user. It's a design bet: some people will pay eight dollars to feel anticipation again.

Milo: From waiting to whimsy: Paw-Paw is a free desktop mascot for Mac that reacts to your typing and clicking. XP unlocks dozens of mascots and over a hundred items. There's one extra purchase at $2.99.

Mia: It's a Tamagotchi for your desktop, essentially, monetized with a single small add-on instead of a subscription. A nice palate cleanser in a week full of agents.

Milo: Then the creative heavyweight of this group: Tonefold. It's an open-source AI music composer, GPL licensed. It generates editable, layered MIDI via Claude or Ollama — so you can use Anthropic's model or run local models. It exports MIDI, WAV, and stems, and it has a producer mode where you approve things layer by layer.

Mia: Two things make Tonefold interesting versus the flood of "AI writes you a song" apps. One: it outputs MIDI you can edit, not a finished audio blob — the human stays in the loop as the actual composer. Two: local models via Ollama means your music creation doesn't have to depend on a cloud. And the producer mode with approvals — there's that propose-and-approve pattern again, now in music.

Milo: And for language learning: ChikyTutor. A voice-based tutor with an interactive whiteboard, in over 70 languages. They've partnered with L'école de français for a hybrid French offering — AI plus an actual school — and it just launched on Android.

Mia: The partnership detail is the notable part. Rather than replacing teachers, they're pairing the AI tutor with a real French school — hybrid learning. That's a more grounded go-to-market than "AI replaces your tutor." The whiteboard plus voice combo makes sense too, because language learning is multimodal — you need to see the script while hearing it.

Milo: Alright. Before we get to the fun stuff at the end, there's one more theme worth pulling out on its own, because it's arguably the design frontier for everything we've discussed: judgment and human oversight.

Mia: The purest expression is Judged.systems. It's an API for judgment in support workflows: a ticket goes in, and it comes back as a choice with probabilities. You can get a score or a probability, and you set thresholds — above the line, the system accepts the decision; below it, it goes to a human for review.

Milo: That's the entire propose-and-approve philosophy, compressed into an API. The threshold is the knob where the organization expresses its risk tolerance. High stakes? Set the threshold so more things route to humans. Low stakes? Let the machine accept more. It's honest about uncertainty in a way a lot of AI products aren't.

Mia: And notice how the other tools in our episode implement the same idea in different shapes. Clippo gives humans a visual canvas to steer agent work. Leanback keeps engineering managers in the loop with daily check-ins rather than acting invisibly. Even Polylane stops at the pull request. BotBus's whole product is approvals. Nobody in this crop is saying "full autonomy, no oversight.

Mia: " The consensus — well, not consensus, but the pattern we can observe across these launches — is that trust is the bottleneck, so human review is being designed in, not bolted on.

Milo: And the honest open question for all of it: as the tools prove reliable, will the humans actually stay in the loop, or will approval fatigue set in and everyone just clicks yes? That's the thing to watch over the next year.

Mia: Okay. Deep breath. Let's end with hardware, safety, and fun — proof that not everything this week was an agent.

Milo: Starting with the strangest: Moonwalkers Dusk. These are motorized shoes that take you up to 7 miles per hour. This version is about 40 percent lighter than before, comes in three sizes, and costs $1,199. The first batch is 200 pairs, shipping October 12.

Mia: Two hundred pairs, so scarcity is real, and the price puts it firmly in early-adopter territory. The weight reduction is the meaningful spec — motorized shoes live or die on whether you'd actually wear them around. Whether the 7 mph experience is pleasant, though, that's something no launch copy can tell you.

Milo: From speed to safety: Aegis is a personal safety app. SOS connected to a 24/7 dispatch service, live GPS, audio and video, a medical profile, a duress PIN, and protective contacts.

Mia: The duress PIN is the detail that shows real design thought here. If you're being coerced, you enter a secondary PIN that looks like it disabled the alarm but actually silently escalates. Live GPS plus audio/video plus actual human dispatch means it's not just a panic button — there's someone on the other end. It's a serious product for a serious need.

Milo: Then three lighter ones. Drunken Penguins: party card games in the browser. Rooms with a code, secret votes, no app to install. The starter deck is free, expansion packs are $1.50.

Mia: No-app browser games are the right call for party games — the friction of "everyone download this" kills the vibe. Secret voting is the mechanic that makes those games work. And $1.50 packs is a very gentle price point.

Milo: Markdoc is a collaborative Markdown editor — source and preview side by side, comments, suggestions, and publishing straight to a GitHub gist or repository.

Mia: The GitHub publishing is the differentiator. If your workflow ends in a repo anyway, editing in a proper collaborative editor and shipping to a gist without leaving the tool is a real shortcut, especially for docs and technical writing.

Milo: And TractionWave: predicts attention on ads with AI heatmaps, working inside Canva and Figma. One free evaluation, then credits — no subscription.

Mia: The credit model fits this use case well. You don't need attention prediction every day; you need it when you're finalizing a campaign. A subscription would be wrong for that. The caveat, as with any attention-prediction tool, is that it's a model of where people look, not proof of what converts — worth treating as directional input on your designs.

Milo: So there it is. Agents getting real management infrastructure, agents doing real work in terminals and offices, models getting small and cheap and almost too human, routing and self-hosting maturing, personal tools with personality, judgment and oversight as the design frontier —

Mia: — and motorized shoes. Thank you for spending twenty minutes with us, I'm Mia.

Milo: And I'm Milo. We'll see you next week.