
1008 | Agentic Everything: This Week in Launches
Show notes
A fast tour of this week's launches: Europe's biggest open-weight model and Google's new image editor, a wave of agents that actually finish work (from video editing to accessibility audits), privacy-first apps from Germany and Finland, and a batch of small developer tools worth knowing about.
Timeline
- 00:00:04 Opening
- 00:00:49 Frontier and open-weight models
- 00:03:30 Agents that do real work
- 00:06:28 Agents for business processes
- 00:10:43 Consumer, care, and private-by-default apps
- 00:14:35 Native, private developer tools
- 00:18:11 Privacy infrastructure and ecosystem watch
- 00:20:56 Closing
Related links
- Mistral Large 4
- Nano Banana 2.1
- Figma Agent
- Introducing Video Editor in Manus 2.0
- Pegasus 1.6 by TwelveLabs
- DevAlly AI Agent
- Ana by Vertice
- Proofsource
- GenPage 3.0
- Ownfeed
- Knuff App
- Rhem Labs
- IrisGo for Solopreneurs
- Supademo AI Demo Agent
- ParakeetAI 2.0
- Velozity
- Redlamp
- Cinch – Reminders & Tasks
- Thalia
- Kitbar
- Databench by Alkera
- Rool
- Vresk
- Temp Mail
- Plugins Radar
This episode is produced by Bri. Bri uses advanced AI technology to turn the feeds you care about into podcasts made for listening. Contact us at hi@bri.so.
Transcript
Mia: Welcome back to the show, everyone. I'm Mia, and with me as always is Milo. Milo, there's a thread running through everything we're covering today, and I think it's worth naming right up front: it's the tension between letting AI do more work for you, and keeping control over what it does and where your data goes. Almost every story we have connects to that in one way or another.
Milo: That's exactly right, Mia. We've got big model news, agents that produce actual deliverables, agents working inside business processes, consumer apps built around privacy, small developer tools, and then a look at the infrastructure underneath all of it. And throughout, the same question keeps coming back: who has the final say? Let's start at the foundation layer, because that's what everything else is built on.
Mia: So the headline is Mistral Large 4. This is a trillion-parameter model, but it's a mixture of experts, which means only 49 billion of those parameters are actually active for any given token. That's the trick of the MoE architecture: you get the capacity of a trillion-parameter model without paying the compute cost of running all of it. And the context window is a million tokens, which puts it in the frontier tier.
Milo: But the part that really matters here is the weights. Mistral says the weights are coming at the end of October. This isn't just an announcement of a closed API model. If they actually ship open weights at this scale, that's being pitched as the strongest open-weight model out of Europe or the US. That changes the calculus for any team that wants to self-host.
Mia: Exactly. Think about what that means in practice. Right now, if you're a company with real privacy or latency requirements, running something at the frontier locally has been mostly out of reach. A million-token context window, on your own hardware, with weights you control — that's the promise being made here. It shifts what self-hosted teams can even consider running.
Milo: Now, we should be careful about what we know versus what we don't. The parameter counts and context length are what Mistral is claiming. Real benchmark performance — how it actually stacks up against the closed frontier models in practice — we don't know yet. And the licensing details aren't clear either. Open weights can still come with usage restrictions, so until we see the actual license at the end of October, teams shouldn't assume they can do anything they want with it.
Mia: Good caveat. And alongside that, there's Google's latest image model, Nano Banana 2.1. This one adds mask editing, which is a genuinely practical feature — you can tell it to change a specific region rather than regenerating the whole image and hoping. They're also claiming better consistency, which is the perennial problem with image models: keeping characters or objects coherent across generations.
Milo: The output range is 1K to 4K, so you can work at higher resolutions directly, and there's Search Grounding — meaning the model can pull in search results to inform what it generates. So on the one hand you have Mistral pushing the frontier of what can be open, and on the other, Google pushing capabilities into a closed image model. Both are foundation layers.
Mia: Which is the natural segue — because a foundation model is just a starting point. What's interesting is what people build on top. And this next batch is agents that produce real artifacts, not just chat.
Milo: Let's start with Figma Agent. What sets it apart is that it works directly with real components and real files on the canvas. That's a meaningful difference from tools that generate a mockup in a separate window and make you rebuild it. Here, the agent is operating inside the actual design file, touching the actual design system.
Mia: And there's bulk editing, which is the unglamorous but hugely valuable part. Anyone who's had to rename or restyle something across dozens of frames knows why that matters. Plus, there's a skills system you invoke with a slash command — so you can define how the agent works and trigger it inline, similar to how developers use slash commands in other tools.
Milo: The human still keeps the final say, though — you're approving what happens to your files. And that pattern shows up again in Manus 2.0's Video Editor. This is an agent that takes on the whole production pipeline: it does the research, plans the shots, picks music, even generates code-based graphics. Then the editor hands the final cut back to the user.
Mia: That last step is the important one. The agent assembles everything, but the end cut is yours to adjust. It's not "here's a video, take it or leave it" — it's "here's a draft with all the pieces done, now you make the call." That's the same philosophy as the Figma agent: automate the labor, keep the judgment with the human.
Milo: And then there's a third one that's from a completely different domain: Pegasus 1.6. This is about robotics. It takes egocentric video — first-person footage, including teleoperation footage — and turns it into labeled, timestamped training data for robots, with no manual labeling.
Mia: That's a big deal because robot training data has always been expensive. Someone has to sit down and label what's happening in every clip. If you can take video that's already being captured — someone wearing a camera or teleoperating a robot — and have it converted into structured training data automatically, you're removing one of the biggest bottlenecks in robotics development.
Milo: Right. So three agents, three domains — design, video, robotics — but one common thread: they produce real artifacts, and the human keeps oversight. Now, the next segment is also agents, but in a different setting. These are agents working inside business processes — compliance, procurement, marketing.
Mia: Start with DevAlly Agent. It records user journeys in plain text — so you describe a flow, like a checkout or a signup, the way you'd explain it to a colleague — and then it checks every step against WCAG, the web accessibility standards, and proposes fixes.
Milo: Accessibility is one of those areas where the bottleneck has never been knowledge, it's coverage. Teams know they should be auditing, but it's tedious and it slips. An agent that walks the journey step by step and flags violations, then suggests concrete fixes, changes the economics of doing it at all.
Mia: Now the procurement side, and this one is my favorite of the batch: Ana, from a company called Vertice. She's an AI negotiation agent for software purchases. And the numbers behind her are the interesting part — trained on 75 billion dollars of spend, 32,000 vendors, and over two million price points. The claimed average saving is 18 percent.
Milo: The insight here is that software pricing is opaque. Two companies can pay wildly different amounts for the same tool, and most buyers have no visibility into what a fair price is. If you're negotiating against a dataset of millions of price points, you're bringing data to a fight that's usually dominated by vendor information advantage. That's why 18 percent is a plausible claim, even though — and we should stress this — it's the company's claim, not an independently verified result.
Mia: Right, and there's a control mechanism worth noting: Ana only sends emails after human approval. So the agent drafts and prepares the negotiation, but a person signs off before anything goes out the door. That's the same human-in-the-loop pattern we just saw with Figma and Manus, applied to spending money instead of making things.
Milo: Then the go-to-market side. Proofsource — and note, that's Proofsource, not Proofstart — measures daily how often you're mentioned across four AI engines. The idea being that when people ask AI systems for recommendations, what those systems say about your product is becoming a real marketing surface. Proofsource tracks that and lets agents publish and review fixes.
Mia: So it's like SEO monitoring, but for AI answers. And whether "AI mention optimization" becomes a real discipline or not, the mechanism here is monitor-and-act: you see what's being said, and you have agent-driven tools to try to influence it.
Milo: Alongside that, GenPage 3.0 generates AI landing pages per ad, per keyword, per account — and then tests them automatically from A all the way through Z. So instead of one landing page that has to speak to everyone, you get a page matched to exactly what the person clicked, and the testing loop runs without a human scheduling experiments.
Mia: And there's a third go-to-market tool with a very different flavor: Ownfeed. You paste in a URL — your product's URL — and it gives you a feed showing only posts from people who actually need your product. There's a one-dollar test day to try it, and notably, no auto-replies.
Milo: The "no auto-replies" part is telling. This is a tool for finding the right conversations, not for spamming them. It's the opposite end of the automation spectrum from GenPage: instead of automating the output, it automates the filtering and leaves the human to do the talking.
Mia: So across this whole segment — accessibility auditing, procurement negotiation, AI-answer monitoring, landing page testing, and lead discovery — agents are moving into processes that used to require specialists or hours of manual work. And notice how "human approval" keeps appearing. That's going to be a recurring theme as we move into the consumer space, where it gets even more explicit.
Milo: Because in consumer apps, the privacy story is front and center. First one: Knuff. It's a family check-in app built on what they call a "Bubble" model — you share your status with a small group, your bubble. The key claim: all data stays in Germany. No US cloud.
Mia: And that's a deliberate positioning choice. For European families especially, knowing that location and status data isn't crossing the Atlantic is the selling point. The check-in itself is simple — "I'm safe, I'm here" — but where the data lives is the differentiator.
Milo: Then something quite different in tone but adjacent in spirit: Rhem, a table robot for seniors. It tracks vitals, observes conversation patterns, gives reminders — and it has AI agents that can book appointments and make phone calls.
Mia: Which is genuinely ambitious. Care for older adults has a coordination problem: family members are scattered, appointments fall through, medications get missed. A robot that sits on the table and handles the logistics — and actually calls the doctor's office to book something — that's an agent doing real work for a population that often can't manage the apps that would otherwise help.
Milo: And then IrisGo, with a feature called Watch & Learn. You record a task once — you doing it, presumably screen or camera — and it turns that into a reusable workflow. And the data stays local.
Mia: That's a nice consumer-to-power-user bridge. If your aunt shows you how to do something once and it becomes a repeatable workflow, that's teaching captured as automation. And "local data" again — the privacy flag keeps coming up in this whole cluster.
Milo: Then two more that round out the consumer set. Supademo's AI Demo Agent: buyers can interact with a talking demo agent around the clock, at 99 cents per 15-minute session. So instead of scheduling a live demo, a prospect can explore the product and ask questions at 2 a.m. if they want.
Mia: Pricing per session rather than a seat subscription is an interesting model there — you pay for actual buyer engagement. And then ParakeetAI 2.0 for job seekers: a question bank built from real interviews, AI mock interviews to practice against, and live answer help during the actual interview.
Milo: The live part is the one that always raises eyebrows — real-time assistance during an interview. It's the tool's headline feature, though, and the mock interview component is the more defensible side: practicing against questions drawn from actual interviews is genuinely useful preparation. We'll present it as what it is — a claim about capability — and let listeners make their own judgment about the ethics of the live mode.
Mia: Fair. And the last one in this group is Velozity — a multiplayer AI office. Agents share team context, you can bring your own AI, and human approval is required for agent actions.
Milo: "Bring your own AI" is worth pausing on. Rather than being locked into one vendor's agent, you plug in whichever models you prefer, and the office becomes the shared space where agents and humans collaborate with common context. And again — approval required. Four out of these six products have an explicit human-approval or local-data story. That's not a coincidence; it's the design response to a real concern.
Mia: Which brings us neatly to the developer tools segment, where privacy and native design become the whole product. And here we actually see the theme show up in small utilities rather than big platforms.
Milo: Start with Redlamp. It's a native open-source RAW photo editor for Apple Silicon, built in Swift and Metal. It follows a Lightroom-style workflow, it's free, and it's under the MPL-2.0 license.
Mia: Native for Apple Silicon is the key phrase — Metal means it's using the GPU directly rather than wrapping some cross-platform framework. RAW editing is compute-heavy, so running the Lightroom-style develop workflow natively on a Mac, free and open source, is a real alternative for photographers who don't want a subscription.
Milo: Then Cinch, which is a native client for Apple Reminders across iOS, macOS, watchOS, and visionOS. Zero telemetry, iCloud-only sync, and — this is the interesting bit for our audience — a CLI so that LLM agents can interact with your reminders.
Mia: So it's a personal app that's also agent-ready, but on the user's terms. You're not handing your tasks to a third-party cloud; an agent you control can read and write reminders through the command line, and sync happens through iCloud, not some startup's database.
Milo: Then Thalia, which is the security-conscious counterpart for coding agents. It's a free native SwiftUI Mac app that wraps Meta's Muse Code CLI. And the feature list is all about control: forced plan mode, so the agent has to present a plan before acting; command approval, so nothing runs without your OK; snapshot review; checkpoints; and no telemetry.
Mia: That's basically a safety harness for a coding agent. Plan-then-approve-then-checkpoint is the pattern people have converged on for trusting autonomous coding tools, and Thalia just makes that pattern mandatory rather than optional. Free, too.
Milo: Kitbar next — a native status bar app for Mac and Windows that puts Vercel, Cloudflare, GitHub, GitLab, Stripe, Polar, Claude Code, and Codex in one glanceable place. No account required, pay once. That's the old-school indie model: buy it, it works, no subscription, no sign-up feeding someone's analytics.
Mia: And rounding out the segment, Databench — an open-source multiplayer workspace under Apache-2.0, with reactive notebooks where agents work alongside the team, and results that stay traceable.
Milo: Traceable results is the privacy-adjacent point there: when an agent contributes to a notebook, you can see what it did. In a team setting, that's what makes agents trustworthy enough to share a workspace with.
Mia: So the pattern across Redlamp, Cinch, Thalia, Kitbar, and Databench: small, privacy-respecting tools that slot into workflows you already have, rather than asking you to migrate to a new platform. Native apps, open licenses where it counts, no telemetry, no accounts. And two of the products we mentioned earlier — Cinch and its iCloud approach, and tools like it — connect directly to what's coming next, because that privacy-first instinct is also showing up at the infrastructure level.
Milo: Yes — the foundation under all of this. Rool is a private AI machine hosted in the EU, specifically in Finland. It runs its own models, doesn't use your data for training, and has a free tier to get started.
Mia: That's essentially sovereign inference. If the Mistral story earlier was about open weights letting you self-host, Rool is the managed version of the same impulse: run AI inside the EU's legal jurisdiction, with a contractual guarantee that your conversations aren't becoming training data.
Milo: And there's a companion product worth mentioning alongside it: Vresk. One workspace for open models, where the system automatically picks the right model per task. Chats never train models, and it's 20 dollars a month.
Mia: So Rool answers "where does my inference run" and Vresk answers "which model should I use" — auto-selection across open models, with the same no-training guarantee. Together they sketch a picture: you don't have to choose between capable AI and control, at least according to these companies' own descriptions.
Milo: And then the last piece of the segment is about watching the ecosystem itself. Temp Mail — the free disposable email service that needs no sign-up — has added private forwarding aliases on Android, OTP copying, and support for custom domains, across 29 languages.
Mia: Disposable email plus forwarding aliases is a practical privacy stack: you give every service a different address, and if one leaks, you kill the alias. It's the small-utility version of the same philosophy.
Milo: And then Plugins Radar, which is the most interesting data point in this whole section. It tracks the ChatGPT plugin directory — over 5,000 plugins — and it's been running 192 searches. And here's the finding: three out of four of the top-10 results have the search term in their name or their tagline.
Mia: Which is a classic SEO pattern showing up in an AI marketplace. When ranking rewards keyword-stuffed names rather than actual quality, discovery gets gamed. Plugins Radar also sends alerts when rivals appear, so builders can monitor their niche. Whether it's a real systemic problem or an early snapshot, that's genuinely an open question — how discovery in these directories evolves is unknown, and it matters a lot for anyone shipping plugins.
Milo: So let's pull the threads together. We started with foundation models — Mistral's trillion-parameter MoE with open weights promised by end of October, and Google's Nano Banana 2.1 on the image side. Then agents producing real artifacts in design, video, and robotics, always with human sign-off. Then agents in business processes — accessibility, procurement, marketing — where the same approval pattern recurs.
Mia: Then consumer apps where local data is the headline feature, then small native developer tools with no telemetry, and finally the infrastructure layer: EU-hosted inference, disposable identity, and a tool that watches the ecosystem itself. From trillion-parameter weights down to a status bar app, the question is the same: how much do you delegate, and how much do you keep?
Milo: And the honest open questions are worth repeating: Mistral's real benchmarks and license terms are still unknown, Ana's 18 percent savings is a vendor claim, and whether AI plugin marketplaces can resist SEO gaming is unresolved. We'll keep watching all of it.
Mia: Thanks for listening, everyone — we'll see you next time.