← Archive Podcast Summarizer
Source ↗
202608.13

2026-08-13 · 00:28:41 · MTS · Media

Are Agents About to Replace Software Engineering?

Are Agents About to Replace Software Engineering?

Abstract

Lauren Tan and Roshan Sadanani discuss Grokbot and Grok 4.6 on MTS Live, the day Cursor and SpaceX announced the model. The conversation treats Grokbot as a persistent colleague with an identity and its own computer rather than a coding panel, and they drive it live on air. Roshan describes internal product-market fit across functions and a chief-of-staff that coordinates shopping, posting, and feedback bots. Lauren traces the idea to Benny, a dog-avatar bot that fixed bug reports overnight, and argues that software engineering is becoming management of agent teams. They close on efficiency, taste, and Grok 4.6’s cost-to-intelligence trade-off.

Prose edition

The show opens on a stray piece of origin story. Lauren Tan has been talking about Benny, a little dog-avatar bot she built because she is, by her own account, fairly lazy, and Cursor was drowning in bug reports. She wanted the reports fixed while she slept. That experiment, more than any slide, is how Grokbot starts to make sense: not as a coding panel with extra knobs, but as a thing you give an identity, a computer, and a night shift.

Then the room resets. An MTS host welcomes Lauren and Roshan Sadanani from Cursor. Earlier that day Cursor and SpaceX had announced Grok 4.6. The plan is to talk about the model, and also to drive Grokbot live, on air.

What the new model changes

What is it, the host wants to know, and what does the new model change?

Roshan treats the question as a product one. Grokbot had gone out the day before and the reception was loud. The motivation was simpler than the stack: make it easier to live with agents. Internally it had the kind of product-market fit teams tell stories about later. Go-to-market, ops, engineering, product — people just latched on. It talked to the tools they already used, and it had its own computer, so the work did not stop when they stood up from theirs.

The host asks for the moment it tipped. Roshan’s example is almost too small. He used it to book Odyssey IMAX tickets. Seventy-inch screen. The host laughs; a new benchmark. The serious version of the same story is persistence. Grokbot keeps going if you are not at the machine.

That is roughly how Cursor builds, Roshan says. Ship internally. Their colleagues have high taste for agent products. Look for the internal fit, then let it out. What was unusual here was the indifference to role. Someone in every function found a use, which is what they want to show.

The host has heard enough preamble. Take it away.

A team of bots

Roshan starts on his own team of bots. A coordinator sits at the center, a chief of staff that knows his calendar and can message the others. It is checking in with Shoppie, the one that buys things on the internet, asking about recent purchases. Tickets. Groceries. That last one had been a favorite internal joke that became a real workflow.

The agents are signed in. Just before the show he stood up a marketing bot on his LinkedIn. He opens its computer: Chrome, logged in, recent posts already scraped off the page. He asks it to post that they are talking about Grokbot live. The host adds, don’t forget to tag the right MTS account. MTS Live.

The design bet, Roshan says while it works, is that this should feel like chatting with a colleague, not waiting your turn behind a thinking spinner. The bot starts trying to post.

He has another one going. Internally there is a channel called Grok Bot wins. He asks the chief of staff to pull feedback from it and put it into Figma Slides, something they could actually share. Let them cook.

The LinkedIn post is already moving. The host can see the computer. It is fast. It finds the MTS Live account. It is the correct one. They will clip this later, the host says, so the shout-out lands twice.

Identity and Grok 4.6

The release was yesterday. What was the reception, maybe from Lauren’s side?

People love the identity, she says. The animation, the logo. It is fun, and the use cases are intuitive. It looks like iMessage. She has seen people hand Grokbot to their mom, and the mom onboarded. There are bugs. Overall it has been very positive, and she is glad they finally got to launch.

And the other announcement, today, next to that one. What is now possible?

The captions wobble on version numbers, but the claim is about Grok 4.6: more intelligent, same token cost, fast enough that Lauren still reaches for it on engineering work. She uses other models too. Grok is the one that is cheap and quick.

Roshan puts a finer point on it. When they launched Grok 4.5 with the SpaceX AI team, the idea was to train models for more than code. Grok 4.6 and the Grokbot investment are the same thesis continuing. They want to stay in the conversation at the frontier, and to bring that into a product someone can actually talk to.

The host has a convert’s angle. Cursor was the gateway into vibe coding, the first tool, still in daily use. Testing Grokbot made the rest of the map obvious: the same system can research, shop, write. The industry still talks as if agents mean coding agents, because the infrastructure for everyday agents has been miserable. OpenClaw was maybe the first mainstream attempt. The host’s mother would never use OpenClaw. The host set one up in Telegram, shopping agent and all, and it breaks, and the job becomes fixing the fixer. Grokbot’s UI is the thing that does not feel like that.

Keep it simple

How did they design it?

Roshan’s brief, a little cheeky and also sincere: keep it simple. A lot of AI software hands you power at the price of complexity. They started with a messaging interface and added the design touches later.

Lauren will not take credit for building it all. The team wanted something clean, because Cursor itself is a power-user instrument, knobs and whistles, tailored for coding. Grokbot needed a different picture: an agent that is a colleague. You do not micromanage a colleague’s tool calls. You talk to it like a person. You can glance at its computer. Elegant, simple.

The host likes having both of them, product and engineering, and then pretends to discover there is a whole team behind them. Lauren plays along. Yes. More people than these two.

Who is it for, if the user is not the typical Cursor user?

Everyone, Roshan says. There is a lot of work sitting around waiting to be made less stupid. The internal motivation was almost anthropological. People at Cursor were already building their own agents, taping chats together, standing up Slack bots. Lauren had a famous one that still lives in the company Slack. That takes a lot of context. The idea was to package it. Those bots were for deployed engineering, for software, for ops. Could they productize that, and make it fun.

Lauren comes back to Benny. The dog avatar. Bug reports arriving faster than any human queue. How do you make agents fix them overnight. That is the persistent-bot experiment. Then people started asking how she built Benny, and whether they could have one, and the thought became: what if everybody could define a bot, give it an identity, a computer, routines. She has bots now that pick up Grokbot feedback, go fix it, and leave her the review.

The host asks the obvious joke. What’s your job now.

It has become interesting, Lauren says. She is almost a manager for digital colleagues. A lot of Grokbot was built with Grokbot. The agents self-triage, they write code, they can kick off cloud agents. She has a demo for that too.

The host loves this rhythm: explain, then just show it.

A spray of messages

After a sponsor break they are back on Lauren’s screen. She has already asked the chief of staff if anything interesting broke today. It found a few things. X is open on the bot’s computer, just in case. She has a Grokbot team, each member responsible for a different feature. The chief of staff is dispatching. In a minute there will be a spray of messages to the teammates, and they will go work. At the end she gets a pull request, videos, a screenshot.

The part that is easy to miss: if you have cloud agents set up in Cursor, Grokbot can launch them. She has one running already, its own computer, a pile of work in flight, a stack of PRs to review later. Software development starts to look like talking to a bunch of agents that do specific jobs, then stepping back and managing them. The old talk about that future, she says, is starting to feel true.

The host notices the avatars. A bunch of cat photos.

Lauren loves cats. She gave them silly profile pictures before the default Grokbot avatars shipped. Those defaults are cute too. This is just having a bit of fun.

Thought to existence

Roshan had hinted at the same shift: more code automated, the human as the person holding the strings. Why is he excited about that, at a human level, for people who do not live in this industry.

One word, he says. Efficiency. There are ideas stuck in people’s heads regardless of role. He has product ideas he wants in Cursor. The limiting factor is the idea-to-existence pipeline. The promise is to throw fleets at the problem and get work back that has taste, that aligns with what he would have put into the product himself. Thought to existence. Compress the cycle. If you look at how Cursor has built itself, that is the direction: more agent tools, more of that complexity, on purpose.

And the software engineering team, from when he started to now.

Lauren answers as someone inside the work. The role becomes engineering manager for your agents. A big part of the job is keeping the codebase in a condition where product managers and designers can contribute high-quality code. She talks about this on Twitter. Invest in refactoring. Rewrite. Encode style as lint rules and CI, hard constraints, so even a not-very-smart agent can thrive. The unlock is leverage. She can empower people to do things they could not do, and she can do things she could not do. She has run large refactors and migrations at Cursor mostly by herself, with one other person.

The host: the cat agents going wild.

Lauren: yes. Work that would have taken a team months, or years.

From the outside, the host says, you might think the demand for engineers falls. On the ground it is the opposite. The thought-to-execution timeline condenses, so something like Grokbot can start as an internal toy, then Grokbot builds Grokbot, then it ships. Is that the trend.

Time-to-prototype is down, Roshan says. The Grokbot team moved very quickly because the tools got better and the models got better. The motivation on the Grokbot side is still: make the assistants more helpful. People still have a strong role. Taste, and whatever else that is. Ideas are not the scarce thing. Seeing more of them become real is.

A Michelin kitchen

Lauren has a metaphor she likes, because she likes food. She thinks of her agents as a Michelin kitchen. She is not cooking every plate. She has hired a staff, trained them with skills and refactors to write code the way she would have. She is the head chef. She assembles the plate. She owns quality. If you invest early in a capable team of agents, you can do much more. Give every engineer one of those teams and you solve more problems, ship more features. Not all of them should be built. The velocity is the point.

The host asks for a hint of what is being cooked internally. Lauren and Roshan are careful, and also not empty. They want Grokbot better. It launched in beta yesterday. Feedback is arriving in every shape. The goal is helpfulness. They are excited about the models, about being in that conversation, about frontier models in the product. Some more fun things later. Bullish on coding agents and on general-purpose ones.

Lauren, as an engineer, wants performance and reliability. She also hacked iMessage support into an early build: text your bots. She does not know if it will ship. People ask for it. The mental model is colleagues. You do not only reach a colleague on Slack. You text them. You call them. Always-on, from anywhere, doing good work.

The host cannot wait to text them, or call them. There is a meme about agents in standup: sorry, I was off all weekend. You are an agent. Maybe they need PTO. Turn the computer off.

Cost and the last pitch

One last thing, because the numbers are sitting there. Grok 4.6 on Cursor Bench 3.2. Grok 4.6 extra high: 70.8 percent, $2.81 a task. Fable 5 Max: 70.5 percent, $17.32. How does a model like that change how ordinary people use agents.

Lauren lives with unlimited tokens and still thinks about this, because users do not. Leaders are having a real conversation about the rising cost of AI. A model as good as the expensive ones at a fraction of the price is the interesting choice. Every company will pick its own mix. If you want cheaper without giving up much, Grok is a serious option.

Roshan’s version: if you have twenty dollars, and the cost per task falls, you just do more tasks. More bots. More code. The industry is playing an efficiency game, and it is a fun one. Intelligence still matters. The benchmark results are cool.

The host takes that to startups and indie hackers who cannot pay token costs out of pocket. The field levels. Then thanks them, and they point at the product one more time. If you have not tried Grokbot, give it a shot.

The agents, somewhere off camera, are still cooking.

SOURCE · MTS · YouTube episode A63sedG-p5Q ↗ · published 2026-08-13 · duration 00:28:41
Speakers · MTS Host · Lauren Tan · Roshan Sadanani
Processing provenance · source notes · English original · provided prose edition