1011 | The Desktop Gets Busy: Agents, Pets, and Plain Files

||Download

Show notes

From baby-proofing a Mac to agent dashboards and Microsoft's rebuilt Windows Search, this episode tours a wave of personal tools, AI developer platforms, and clever desktop utilities — and asks what it means when your computer starts running itself.

Timeline

  • 00:00:04 Opening
  • 00:00:50 Agent workspaces: terminals, canvases and review
  • 00:05:41 Grounding and measuring AI: papers, dashboards, visibility
  • 00:09:52 Creating and consuming content with agents
  • 00:12:33 Desktops and browsers: Linux up, Windows and Firefox refresh
  • 00:15:12 Small Mac utilities: focus, pets, safety, sound, history
  • 00:19:14 Builders' tools: selling, showing, watching
  • 00:21:35 Closing

Related links

This episode is produced by Bri. Bri uses advanced AI technology to turn the feeds you care about into podcasts made for listening. Contact us at hi@bri.so.

Transcript

Mia: Welcome back to the show. I'm Mia.

Milo: And I'm Milo. Today is basically a launch briefing — we've got a stack of new products and updates from the last day, and we picked out the ones with the clearest problem, the clearest solution, and actual specifics to show for it.

Mia: And the thread that connects almost everything here is agents. Agents running in terminals, agents that need grounding, agents making your emails, and the desktops and little utilities that have to live around all of that. So we'll walk through it as one flowing story — from the agent workspace itself, out to the plumbing that feeds agents, through content creation, and finally onto the desktops and the tiny tools on top.

Milo: Right. So let's start where the work actually happens — the workspace.

Mia: Two very different Mac products came out with the same underlying idea: the canvas is the workspace. The first one is put·here. It's a Mac canvas app where you drop in notes, links, and files — by paste, by drag, or with a keyboard shortcut, Control-Option-P. And here's the part that makes it different from the average pretty whiteboard: each canvas is actually a real folder of plain files underneath.

Milo: That's the key detail. So it's not a proprietary database that holds your stuff hostage. If you like the visual, spatial way of organizing things, you get that — but everything lands on disk as ordinary files, which means it plays nicely with everything else on your machine. That matters especially now, because one of the things people do with folders of files is point tools at them, including agents.

Mia: Who's it for? Anyone who thinks in spatial layouts — researchers, designers, people collecting fragments — and it's free to start, so the barrier to trying it is basically zero. The open question is the one we keep coming back to: how does the plain-file and canvas model scale when you've got lots of stuff, or lots of agents poking at the same canvas? We don't know yet.

Milo: Now flip the canvas idea around and you get Hypervibe. Same "infinite board" metaphor, but instead of notes and files, what lands on the board are terminals. It's a native Mac app that runs Claude Code, Codex, Gemini, and over thirty CLI agents side by side, as real PTY terminals.

Mia: Real PTY terminals is worth spelling out. A PTY is a proper pseudo-terminal — meaning these aren't simulated chat windows or wrapped web views. The agents genuinely run in their own terminal sessions, so you can have several of them going at once, scattered across an infinite board, and watch them all.

Milo: Who's that for? Developers running multiple coding agents — which is rapidly becoming a normal way to work. You fire off one task here, another there, and instead of tab-switching through terminal windows, you see the whole fleet at a glance. Versus just opening a bunch of terminal tabs, you get the spatial overview, which is genuinely different when you're juggling three or four concurrent agents.

Mia: The limitation is obvious though: it's Mac-only, and we don't have any evidence yet about how this scales — eight agents screaming output at once on one board? That's the unknown. And that unknown leads us neatly to the next piece, because if agents are making changes to your code, someone has to review those changes.

Milo: That's GitGlow. It's a local Git workspace built specifically for reviewing what agents changed. You get staging, you get history, and the terminal is checkout-aware — so it knows what branch and state you're looking at. The Pro tier adds stale-review marks, which help you notice when a review has been sitting unaddressed. It starts free.

Mia: Why this matters: in the classic workflow, a human writes a diff and another human reviews it. Now the diff is written by an agent, and the review step is the missing link in the whole pipeline. GitGlow is betting that reviewing agent output becomes its own distinct activity, with its own tooling — local, not cloud-based, which fits the review-before-you-trust mindset.

Milo: And there's one more product that pushes the agent trend past the desktop entirely. API Claws, by a company called Buda, is a hosted Agent API. The pitch: if you're building an app or a device, you can give it a cloud agent without running any infrastructure yourself. It includes Drive memory — so the agent remembers things across sessions — model routing, so you're not locked to one model, sessions, and embeds.

Mia: So on one end of the spectrum, Hypervibe puts agents in terminals on your own Mac; on the other end, API Claws says you don't need to run anything at all, the agent lives in the cloud. Those are two genuinely different answers to the same question — where do agents live? — and it's too early to say which bet wins. What both tell us is that agent workflows are becoming the interface itself, not a feature bolted onto an app.

Milo: But an agent is only as good as what it can see and know. Which brings us to the plumbing underneath — grounding, measurement, visibility.

Mia: Start with research. Lune is a search engine and MCP connector that grounds AI agents in full-text papers from top-tier conferences — not summaries, full text — with citation tracing, so you can follow where a claim actually comes from. The evidence offered: it's used by over a thousand researchers, and it's free to start. As always with maker claims, treat the adoption number as a claim, not a verified result — but the mechanism itself is concrete and specific.

Milo: Why it matters is simple: agents hallucinate when their inputs are thin. If your coding agent or writing agent is going to make claims, pointing it at actual peer-reviewed conference papers with traceable citations is one of the more serious ways to ground it. And the MCP connector part matters because MCP is becoming the standard way agents reach out to tools.

Mia: From grounding inputs, jump to grounding outputs — data. Claude Dashboards is in beta and paid, and it turns plain-language questions into live dashboards on your warehouse — BigQuery, Snowflake, and others. There's also a feature called Motion that animates explainers as editable code, with MP4 export. So the same plain-language-to-artifact idea, applied to business data instead of papers.

Milo: Worth being careful here: it's beta, it's paid, and we don't have third-party evidence of how well it performs. The interesting conceptual point is that dashboards — traditionally a thing you build deliberately in BI tools — become something you ask for in a sentence.

Mia: Then there's the measurement side nobody talks about enough: how visible is your brand inside the AI answers themselves? ReSO AI is a GEO platform — generative engine optimization — that audits your visibility in ChatGPT, Perplexity, and Google AI Overviews. It diagnoses gaps with over 240 checks, then turns them into prioritized weekly actions.

Milo: That's a new kind of SEO, basically — except the search box is now a chatbot, and there's no crawler to optimize for in the classic sense. The limitation we should flag: as the underlying models change, we genuinely don't know whether these GEO fixes stay fixed. You could optimize for ChatGPT's answers this month and have it shift next month. That's a real open question for the whole category.

Mia: And one product flips the whole direction. Microsoft released a decision model — that's the actual framing — that returns calibrated probabilities for fixed options in a single pass. It's about 35 times faster than GPT-6 Sol, at $0.042 per million input tokens.

Milo: So instead of generating text and hoping the reasoning lands somewhere useful, this model is designed for the specific case where you have a fixed set of options and you want a probability on each — cheaply and fast. Calibrated is the operative word: the probabilities are meant to actually correspond to reality, not just sound confident. At that price point, it's the kind of thing you could call constantly from an application.

Mia: Put Lune, Claude Dashboards, ReSO, and the decision model together and you see the shape of it: grounding, visibility, and cheap decision calls are the plumbing of the agent era. Agents need trusted inputs, outputs need measuring, and decisions need to be fast and cheap. Okay — from the plumbing, let's move to what agents actually make, and how it reaches people.

Milo: Maildun is an agentic email builder for Mac. It has a drag-and-drop editor, AI editing through your own API key — meaning you bring your key, you control the model and the cost — and, notably, a built-in MCP server so Claude Code, Codex, and Cursor can work with it directly. There are workspaces for organizing the work.

Mia: That MCP detail connects right back to what we were saying earlier. The email builder isn't just "has AI features" — it's exposing itself as a tool that coding agents can operate. Your agent in Hypervibe could, in principle, be the thing driving your email production. That's the agentic app pattern taking hold: the app becomes something agents use, not just something humans click.

Milo: But generated content needs checking, and that's where Onepin comes in. It validates AI voiceover. It normalizes prices and dates — because text-to-speech models mangle those constantly — it checks names against a four-million-word dictionary, it scores every line, and here's the genuinely clever part: it can fix a single word without re-rendering the whole track.

Mia: That last capability is the differentiator. Anyone who's produced voiceover knows the pain: one mispronounced name, and you re-render the entire take and hope nothing else shifts. Fixing a single word in place changes the economics of iteration. It's free to start, and again — the maker describes these capabilities; nobody's independently verified the quality. But the problem it targets is concrete and universally felt.

Milo: And at the consumption end of the pipeline: Fluence. It's an Android app that reads PDFs, EPUBs, and documents aloud, entirely on-device, in 31 languages. No account, no server, no subscription — €14.99 once, and it works offline.

Mia: Look at the whole pipeline we just described: Maildun generates with your own API key, Onepin verifies locally line by line, Fluence consumes entirely on-device. Three different stages, one shared philosophy — one-time purchase, local processing, no forced account. Whether on-device models stay competitive with cloud ones is the open question, but the purchasing model is clearly resonating with makers right now.

Milo: And that local-first philosophy lives somewhere very visible: the desktop itself. So let's talk about desktops and browsers — Linux pushing forward, Windows and Firefox refreshing.

Mia: On Linux, Naarchy is a free, open-source, MIT-licensed top-bar island — built in Rust with GTK4, aimed at Hyprland users. It packs a file shelf, clipboard history, music controls, a timer, and a calendar into the top bar. MIT license means anyone can fork it, and Rust/GTK4 means it should be lightweight.

Milo: Then Skymir, which solves a problem Linux users have genuinely lacked good answers for: native GTK Google Drive two-way sync. Docs open via .gdoc shortcuts, deletes go to trash instead of vanishing, and — my favorite detail — it refuses to sync a deletion that would remove more than half your files. That's a guard against the classic horror scenario where a sync bug or a compromised session wipes your Drive. $19 once.

Mia: That trash-safe delete and the half-files guard are exactly the kind of defensive engineering that separates a serious sync client from a naive one. So Linux is getting a richer top bar and a proper Drive client, both in the one-time or free-and-open mold.

Milo: Meanwhile, on the big platforms, it's refresh season. Microsoft rebuilt Windows Search on WinUI 3 — faster, results in a single list, with inline actions like toggling dark mode, muting, or arranging windows right from search. It's rolling out to Windows Insiders.

Mia: And Firefox 157 is, by its own description, the biggest visual refresh in years: new themes and wallpapers, a customizable homepage, vertical tab groups, and compact mode returns. Note "biggest refresh in years" is Mozilla's claim — but the feature list is concrete. Vertical tab groups in particular is the thing people have used extensions for, now native.

Milo: The pattern across all four is the same: fast, inline, status-rich interfaces. Search that acts instead of just links. A top bar that does five jobs. Browsers consolidating tab management. The desktop shell is becoming denser and more immediate everywhere.

Mia: And riding on top of those desktops — the small utilities. This is honestly one of the most fun clusters. Let's start with safety, because Baby Desk solves a problem every parent with a Mac knows: your kid grabs the laptop and starts mashing keys.

Milo: Baby Desk gives you a fullscreen baby playground that swallows every key and click — including Command-Q — and even screenshots. The parent gets out through a hidden key combination plus Touch ID or PIN. $4.99. So the toddler can pound the keyboard to their heart's content and nothing escapes, nothing breaks, nothing gets posted.

Mia: The design question it answers is deceptively hard: how do you make an app a child cannot defeat, while an adult still can? Swallowing Command-Q, blocking screenshots, and gating the exit behind biometrics — that's a thoughtful answer to a real problem for five bucks.

Milo: Then the wellbeing end of the menu bar. Pawse is free and open source, a macOS menu bar app where a 3D pet walks across your screen asking about water, habits, and breaks. There's an optional hard block that locks your screen — so if gentle nudges don't work, the pet gets authoritarian.

Mia: And then Cue, which is my favorite idea in this whole batch: your AirPods become a focus switch. Put them in, and Focus mode triggers, sites get blocked, a timer appears in your notch. Take them out, and a break starts. Free — though it's not notarized, which means macOS will throw up a warning and you should know what you're installing.

Milo: That's ambient hardware being repurposed as an interface. Your AirPods were already a signal — "this person is listening to something" — and Cue just makes the system respond to it. It fits the wellbeing cluster with Pawse: both are reading context and surfacing status and habits without you opening anything.

Mia: Staying in the sound-and-attention neighborhood, TapNoise gives you twenty keyboard sounds, plus a desktop keycap pet named Tappy. And the forward-looking part: agent-done chimes — sounds when Claude Code, Codex, or Cursor finish a task — and knock commands. From $4.99 once.

Milo: The agent-done chime is a small thing with a real purpose: when you have several agents running — back to that Hypervibe board — you can't watch them all. An audio signal when one finishes is exactly the missing feedback loop. Sound design meeting the agent workflow.

Mia: And the last one in this cluster is History Monk, which replaces browser history as we know it. Search in 80 milliseconds, typo-tolerant, over eleven filters, bulk delete by site — and it's one hundred percent local and private. Seven-day free trial.

Milo: The bulk delete by site is the standout: everyone has that one site they've visited two hundred times and would rather erase from history entirely. Native browser history makes that painful. 80ms search with typo tolerance is a claim, but it's the kind of claim that's at least specific enough to be tested by anyone who tries it.

Mia: So the small-utilities story: single problems, one-time prices, privacy as the default rather than a setting. Baby Desk, Pawse, Cue, TapNoise, History Monk — none of them want a subscription, and several of them want your data to never leave your machine. And if Cue is any indication, more status and wellbeing signals will surface from ambient hardware going forward.

Milo: Which leaves the last cluster — the tools for the people building all of the above. Selling, showing, and watching.

Mia: First, selling. Toolaby adds licences, subscriptions, trials, and paywalls to Chrome extensions in one command. Payments go to your own Stripe account — not through the platform — and it costs either 0.5 percent or $25 a month.

Milo: Running payments to your own Stripe is a meaningful difference. Many platforms take a cut and hold your customer relationship; here the billing infrastructure is the product and you keep the account. The pricing itself is honest about the trade: a percentage scales with your revenue, a flat monthly doesn't. One thing we genuinely don't know is how fee stacking works out for tiny products — if Stripe takes its cut and Toolaby takes 0.

Milo: 5 percent, small sellers should do that math for their own numbers.

Mia: Then, showing. Skreno is a browser video studio. It records screen and camera as separate tracks — which is the crucial detail, because separate tracks mean you can edit them independently afterwards instead of burning your face into the corner of the video at record time. Then you share one link with a transcript and an AI summary, and it handles bug reports with logs attached.

Milo: That bug-report feature ties it back to the developer world we started in — recording a reproduction with the logs already gathered. It's the demo-plus-support tool in one.

Mia: And finally, watching. Kitbar is a thin screen-edge bar that shows the status of Vercel, GitHub, Stripe, and Claude Code as rings. No account, no server — your API keys live in your keychain — and you pay once.

Milo: Same philosophy we keep hitting: keys in the keychain, no middleman server, one-time price. And notice Claude Code is in there — status monitoring has expanded to include your agents, right alongside your deployments and your revenue. The indie developer's whole operation, visible as rings on the edge of the screen.

Mia: So here's the arc of today, drawn in one line: agents became the interface — in terminals on boards, in cloud APIs, in review workspaces. They needed grounding, so papers, dashboards, visibility audits, and cheap decision models arrived. They started making things, and the pipeline from generation to verification to consumption went local and one-time-purchase. And around it all, the desktop refreshed — and the little tools on top went after single problems with real answers.

Milo: The open questions to carry with you: whether canvas and plain-file models scale to many agents, whether GEO fixes survive model changes, whether on-device models stay competitive, and — the one we'll be watching — whether agent-native billing becomes standard for indie products. That last one feels like the next episode's subject already.

Mia: Thanks for listening — we'll see you next time.