
Former DeepMind Eng turned Founder: Computer Agents and Token Economics with Simular AI Ang Li
Om avsnittet
In this episode, I speak with Ang Li, co-founder and CEO of Simular AI, about the rise of computer use agents and his vision for computers that can increasingly do work on behalf of humans. We discuss why computer use matters beyond the current wave of API-based agents, particularly for the vast amount of enterprise work still carried out through legacy desktop software that was never designed to be accessed programmatically.
We explore where these agents could have the most immediate impact, from processing invoices and extracting information from unstructured documents to navigating financial, healthcare and other enterprise systems. Ang argues that the goal is not necessarily to remove humans from the workflow, but to shift the balance between execution and judgment, with agents handling repetitive tasks while people remain responsible for decisions and sign-off.
The conversation also gets into the economics and technical challenges of computer use agents. Ang explains why relying on a frontier model for every click can be expensive, slow and difficult to control, and how Simular’s neurosymbolic approach turns repeated workflows into executable playbooks. We discuss the role of smaller and open-weight models, the changing economics of AI agents, and why Ang sees computer use as a way to make automation more accessible beyond the highest-value technical work.
Finally, we look at what this shift could mean for SaaS and the future of work. Ang distinguishes between systems of record and software that primarily serves as a portal or interface, and argues that these categories may face very different levels of disruption from agents. We end with a broader question about human agency: if AI increasingly handles execution, will the more valuable skill become knowing how to identify the right problems to solve in the first place?
The AI Proem Podcast is part of the AI Proem newsletter, which has ~13k followers globally. To learn more about China AI, the business of AI, and how AI is impacting society, please check out the newsletter here and more insightful conversations here.
Chapters
00:00 Introduction to Simular and Its Vision
05:14 The Future of Work and Human-AI Collaboration
10:21 Technological Landscape and Market Demand
15:31 Real-World Applications and Use Cases
20:40 Challenges and Competitive Landscape
29:55 The Dual Nature of Work: Content Creation vs. Execution
33:50 The Competitive Landscape: Frontier Labs vs. Open Weight Models
37:17 The Rise of Computer Use Agents
43:30 Empowering the Average Person: High Agency Through Technology
48:46 Understanding SaaS: Infrastructure vs. Portals
52:32 The Future of AI: High-End Jobs vs. Repetitive Tasks
AI-generated Transcript (for reference only)
Grace Shao (00:00)
Ang, thank you so much for joining us today. Really excited to have you. To start with, could you tell us a bit more about yourself and why you built Simular? What is your long-term vision for the company? And what really brought you here along your academic journey?
Ang Li (00:13)
Yeah, thank you so much, Grace, and thank you for having me here. My name is Ang Li. I’m the co-founder and CEO of Simular. We’ve been working on this company for almost three years.
In short, we call ourselves the autonomous computer company, meaning we are building autonomous computers. The goal of Simular is basically this: everyone has computers right now, but we have to work on them manually by moving the mouse, typing on the keyboard, looking at a screen, and understanding what’s going on ourselves.
We envision a future where computers will do the work on behalf of humans, on their own. That’s really the technology that we’re building towards. Nowadays people also call them computer-use agents. It’s basically a general-purpose agent that can use the computer just like a human.
A bit about my background: I’ve spent almost 20 years researching AI. My personal research direction has basically been trying to figure out what people sometimes call AGI, artificial general intelligence. The goal is to have a general-purpose system that can learn like humans and perform actions just like humans.
I see this autonomous computer problem as the first possible realization of AGI technology that could have a huge impact on society.
Grace Shao (01:29)
It’s quite interesting. You mentioned computer-use agents. It seems like there’s been quite an influx of capital going into this space right now in Silicon Valley. Has there been a genuine technological shift here? Why is everyone suddenly so hyped up about CUAs?
Ang Li (01:44)
Yeah, so that’s the interesting part. We started three years ago, and when I told people we were building agents, people didn’t understand it. I was telling people, “Okay, we’re building agents that use computers.” And people would ask, “Why? Why are you building agents that use computers? Why not just use APIs?”
Agents aren’t actually a new concept. It’s a word that has been around for tens of years in the research community. We already talked about agents within DeepMind when we were doing research towards AGI. It just wasn’t very familiar to the broader public.
Three years ago, we had the first version of ChatGPT, and we knew the scaling laws for foundation models were working. Foundation models were becoming very powerful.
We looked at the trajectory of the technological shift and realized that, for a computer to work like a human, you need APIs. If you want a computer to go into your Gmail, look at an email, and send an email on your behalf to someone else, there’s already a Gmail API. So if you want to do those kinds of tasks, you just let the agent call the Gmail API.
Three years ago, that was basically the case for tool calling: basic APIs.
Then we asked: suppose we have all the APIs readily available to agents, what’s remaining?
The answer became very natural. What’s remaining is all the software that has no APIs.
For example, lots of companies have legacy software on Windows computers, like old ERP systems that nobody is maintaining anymore. That software has been running for many years, and people don’t really want to change because they’re so familiar with it.
The problem is that because the software is outdated, people still have to manually work on it. There are no APIs.
If our goal is to liberate human labor, we don’t want people spending so much time sitting in front of computers doing repetitive, tedious tasks. Nobody likes that. Everyone thinks, “Why can’t I do something more interesting with my life? Why should I spend eight hours every day doing repetitive stuff that doesn’t require much cognitive load, just looking at a spreadsheet and filling out the same form over and over again?”
Those problems cannot be solved by agents that only have APIs.
So then we realized this is actually a harder problem. It requires technology that can look at a screen, decide, “Should I click this button here? Should I type something?”, move the mouse to the coordinates of the button, click on it, and move to the next page.
It feels a bit like a self-driving car in the digital world. A self-driving car looks at the streets and decides, “Should I turn left or right? Should I press the gas?” In this case, we’re looking at a screen and deciding where to click.
When you combine the two, API agents and computer-use agents, you cover the full spectrum of computers. There’s nothing else remaining.
Once you have API agents and computer-use agents and combine the two, computers can become autonomous. That’s the AGI that everyone is striving for.
Grace Shao (05:25)
It’s pretty crazy. Your vision of the future of work is essentially that computers run themselves.
I get that vision. But wouldn’t work itself, by nature, just change? So much of our work right now, like you said, filling out PowerPoints or spreadsheets, you can almost call it performative because it’s ultimately for humans to view.
But if it’s agents or computers viewing the output, do we still need to fill out those PowerPoints and forms?
Ang Li (05:50)
Yeah. I think first we have to look at what the bottleneck of work is right now.
If we still have a lot of people performing manual data entry, that’s the bottleneck right now. If we remove that bottleneck, people’s productivity could be 100x in the future, and they can spend more time on strategic decision-making instead of doing this performative work.
That’s our first goal as a company. Why not just remove that repetitive work from people?
It doesn’t mean that in the future computers do all the work and humans never look at a screen. It’s more like when you hire an intern. You delegate some tasks to the intern and say, “Can you fill out this form?” The intern comes back and says, “I finished it. Do you want to take a look?”
You still take a final look and make sure everything is correct and according to the company’s policies.
Computer agents will do the same. It’s not going to be that agents just finish the work, make some random mistake, and walk away.
In the end, the human will still be the final gatekeeper for everything. Humans just don’t have to be involved throughout the entire process. You still need a human to sign off.
Grace Shao (07:23)
Right. You don’t need to be hitting enter, enter, enter the whole time when Claude keeps prompting you.
Ang Li (07:26)
Yeah, exactly.
Grace Shao (07:30)
But then my question for you is: do you think the future will still look like the desktop we know today?
What would the interface be? Are we still going to use the current desktop applications we use today? Or do you have a different vision for how humans even interact with AI?
Ang Li (07:47)
To answer this question, I think we have to clarify two concepts.
One is what kind of device humans use. The second is what kind of device agents use. Those two devices don’t have to be the same.
First, we have to look at the natural way for humans to receive information. This seems like a relatively simple problem.
Humans invented paper, and society has been using roughly this size of paper to view documents, do approvals, and sign off on things.
For computer screens, we envision the future as something more like an iPad.
You don’t necessarily have to have a keyboard or trackpad. You can just have a big screen.
Grace Shao (08:36)
You don’t even necessarily have to have a screen or interact with it, right? It could just be like a box.
Ang Li (08:42)
Yeah. I mean, you can still have interaction. You can still click on it.
But there’s a fundamental size that humans are comfortable with. It’s roughly the size of a piece of paper.
Some people say the future is the mobile phone. I disagree with that because the size is limited. It’s hard for me to read books or documents on it.
An iPad-sized screen is actually a good size. It’s kind of like current computers, just removing the keyboard and trackpad.
The second question is: what kind of device do agents use?
That device doesn’t necessarily need a screen. It’s just a machine with computational power, and that’s good enough.
Whether the machine is a laptop or a desktop doesn’t really matter. What’s important is the computing power. You could have GPUs in there. You could host models in there. You should think about it more like a piece of metal sitting in a data center. That’s the machine the agents use.
Then there’s another question along that line: what kind of operating system runs on the machine? Is it still desktop, mobile, iPadOS?
Our view is: why not have a common infrastructure for all operating systems?
For now, the most powerful one is the desktop because the majority of the killer applications for computer use are legacy desktop software.
That becomes the primary direction we’re tackling right now.
But it’s possible that in the future, if a lot of apps run on mobile phones, then we need Android machines in the cloud that help you offload that kind of work.
It’s definitely possible. It’s just that right now, we see the biggest bottleneck as Windows desktops.
Grace Shao (10:40)
So in the future, would you guys look at creating the operating system you were just talking about? Or would you even go into the hardware?
Ang Li (10:48)
Yeah. I mean, we are an autonomous computer company, so basically we’re working on computers. It’s definitely possible.
The future trajectory could go beyond software.
But right now, we already have cloud infrastructure where, if you say, “I need a Windows desktop,” we can give you a Windows desktop in the cloud. If you need an Android phone, we can give you an Android phone. If you need a Mac, we can give you a Mac.
We have infrastructure that can allocate any operating system in the cloud for your agents to use.
It’s general-purpose, and the agents can choose how many devices they want.
Some people might want to use 10 Windows desktops for their own purpose. Some companies may need 100 to run massively parallel jobs for certain software.
That’s already possible today.
Grace Shao (11:37)
I want to double-click on something you mentioned earlier.
You sound more optimistic about the idea that, as AI advances, we’ll be freed up to work on more creative work.
I’m going to play devil’s advocate here. Some people would argue not everyone wants to work on strategic work. Not everyone has that creativity, not everyone wants to do that, and frankly, not everyone has the capability to do that.
Some people have been part of the execution chain for the last three decades. That’s how the workforce has trained them.
So it leads me to this broader conversation. There’s a bit of fear-mongering in Silicon Valley saying AI is going to take our jobs. AI will certainly take over many of the administrative and executional tasks you mentioned.
But others are saying jobs don’t equal tasks.
It sounds like what you’re proposing is that CUAs will help with tasks but not necessarily replace the entire job.
Help me understand your more philosophical view on this and how you see the future of work.
Ang Li (12:34)
Yeah. We talk to customers, and people also reach out to us.
I can give you an example. There was a general manager of a car dealership who reached out to us. It’s a small family-style business in the U.S., only three to five people.
They go into QuickBooks and generate hundreds of invoices every day for their customers, and they don’t like it.
Even though people are willing to do this job, they don’t like it. That’s the real problem for society right now.
They earn their wages through this kind of work, but it’s not really a job they like. If those people had the opportunity to do something else, they would do it.
The real problem is not that they’re trained to do this, therefore they want to do it. It’s because this is the way for them to earn wages.
Suppose we had a way to give these people the same amount of salary and the opportunity to do something else they’re interested in. Everyone would do that.
The real question is: do we have enough productivity that allows people to explore their interests?
My view is that this division of labor exists because we are in a constrained economy.
When you only have a certain amount of money and resources to distribute, you have to make trade-offs.
But if society’s productivity becomes 100x higher, meaning we produce 100x more goods and resources with the same population size, then we can allocate more funding and resources to each individual.
In that scenario, people aren’t going to say, “Because I was trained to do repetitive work, that’s what I want to do.”
People will ask, “Can I use Claude Code to create apps?”
A lot of people are already doing that. People in non-technical industries who used to do tedious work are turning to coding agents and asking, “What kind of thing can I create?”
Everyone becomes a creator. You’re creating something new. That’s what people find interesting.
I wouldn’t doubt that.
I feel like the main problem is that our society has a limited economy. That’s waiting for us to amplify it by 100x.
This digital workforce is an opportunity for society to amplify the economy because, in the future, every company could have 100x more digital workers who aren’t human.
Naturally, the speed at which you produce goods or run operational pipelines becomes much faster. Companies run faster and produce more resources for society.
That’s the opportunity I see. This might be a little controversial, but I see agents helping industry become much more productive, and in return giving humans more opportunity to do creative work or whatever work they’re passionate about.
The problem right now is that we are constrained, so people are forced to do a lot of manual work.
Grace Shao (15:56)
No, I actually agree with you on this.
Technology has always disrupted jobs, but that disruption has also led to replacement and redirection of people’s interests.
Even looking at the last generation of workers, think about people who worked in car manufacturing or factory jobs. As automation replaced some of those jobs, people found new work.
What you’re saying is that now our minds are doing these laborious jobs. They’re almost mental labor jobs.
Once our minds, or at least our time, are released from that, potentially we find new ways to use them.
But that may take a decade or two, or even a generation, to figure out what the new way of living is.
I’m in that optimistic camp as well.
It’s just interesting when you talk about 100x supply: will there also be 100x demand? And how might that affect the economy?
But we’re not going into that today. That’s a whole rabbit hole we could go down.
Ang Li (16:54)
Yeah. In short: use the new technology, find new opportunities, and make the pie bigger.
You want to make the pie bigger so everyone can be relieved from some of the stress and burden of labor.
Grace Shao (17:09)
I feel stressed out covering AI right now because it’s moving too fast. There’s too much to follow every day.
But look, what are some real-life use cases you’re seeing?
I liked your dealership example, but what are some more enterprise-facing use cases where people are already using Simular?
I think on another podcast you were talking about healthcare and pharmaceuticals. Are those verticals something you guys are focused on, or have they naturally become high-demand industries for your kind of technology?
Ang Li (17:44)
It’s organic. They naturally become high-demand industries. Financial services as well.
We’re working on this general-purpose horizontal platform that allows people to take our technology, build on top of it, and solve their own problems.
We have this market pull where people come and talk to us.
We’ve found a pattern across almost all the use cases.
The pattern is basically: “I have a form in a certain shape.”
It could be a physical receipt. It could be handwritten or machine-printed. It’s basically unstructured data, often an image.
The agent needs to look at the image, parse the information on that paper, and move the data into another system.
And that system usually doesn’t have APIs.
It’s often a desktop environment. Many of them aren’t browser applications. They’re old desktop applications running on Windows computers.
This kind of manual data entry is one of the biggest bottlenecks in industry today.
And it’s not just healthcare. People from many other industries have exactly the same problem.
Grace Shao (18:55)
So that’s currently the biggest use case.
Help me understand the technicalities of this. Am I giving you access to my PC? How does this work?
Ang Li (19:05)
There are two types of scenarios.
In regulated industries, people worry about giving a third party access to private information.
Suppose you’re a healthcare provider. You already have your own computer with the software installed.
What you want is: “Can you install an agent on my computer so the agent takes the image, fills out the form, clicks through the portal, and does some calculations?”
That’s the option we call BYOD, bring your own device.
Basically, you have your own device. You have everything installed there. I don’t touch any of it. I give you the agent, and you install the agent as software on your computer.
Then there’s another set of users who want to scale.
They say, “Okay, I have one computer. I can make it an autonomous computer. But what if I want to run 100 autonomous computers for this type of task in parallel?”
It’s impossible for them to stack 100 computers in their office and manage all of them.
So they come to us and say, “Can you provision 100 virtual computers in your cloud?”
That becomes the advantage of having cloud infrastructure for all these desktop computers.
In that scenario, you don’t need to provide the computer. You come to us, and there’s already a computer there with an agent installed.
You just prompt the computer and say, “Okay, do this task. Fill out this form and test our computer,” and then you walk away.
You don’t have to manage anything.
Grace Shao (20:40)
I see. So a lot of it is actually outsourced to you guys to manage.
The elephant in the room for all the newer players is obviously Big Tech and the frontier labs: OpenAI, Anthropic, Google and others.
These companies are becoming very good at almost everything. They’re going around eating everyone’s lunch.
What happens when they become very good at computer-use agents?
What is the lasting moat for a company like yours? You’re three years old, you’ve obviously done very well in your vertical and were very early to it, but how do you create a lasting moat?
Ang Li (21:16)
First, I want to thank the frontier labs for working on this problem.
When we started three years ago, I had a really hard time convincing people this technology was useful. People thought it wasn’t useful.
Another group of people would always ask me, “What if OpenAI does the same thing?”
It feels natural because we are working on a very general technology. It feels natural for OpenAI to do the same thing.
And I didn’t really have an answer to that.
I would literally tell them, “Okay, it’s definitely possible.”
I came out of DeepMind and I’m working on this general technology. I believe there are a few sets of people in the world who share the same vision and are pushing the same technology toward the future.
About a year after that, OpenAI started working on this problem. Then other labs started working on it too.
Even though we were one of the only companies focused on it at the time, we released our first open-source computer-use agent, which ranked number one on a public benchmark called OSWorld. That happened in October 2024.
I still remember that one week later, Anthropic released its first computer-use product.
That was the moment where it became real: frontier labs were working on the same problem as us.
The interesting part is that it actually helped us educate the market.
More people became familiar with the technology. Whenever they talked about computer use, they knew there was a company called Simular that had been working on this for years.
They thought, “It seems like promising technology. We want to try it and see how it can help us as a company.”
So that’s the first answer: having other people working on the same thing may not be a bad thing.
I want to say this to founders because everyone worries about a big company taking over.
Actually, it may be the contrary. If a big company starts working on the same problem, it validates the product-market fit for your technology.
Then we have to think about what’s next.
Our philosophy has always been: if there is something other people can work on, why does the world need us to work on it?
In the beginning, we wanted to build the whole computer agent. We knew it was relatively easy for people to build API agents, so why should we spend time on API agents? Why not spend time on the remaining problem?
That’s why we worked on computer-use agents.
Now the big labs are working on computer-use agents as well. So we have to look at what kind of technology they’re focusing on.
The big labs’ business model is based on the idea that AGI is a single model.
Their business model is to sell APIs for that single model.
Computers are a vertical on top of their model business. Computers are not the whole thing for frontier labs.
They’re still trying to train the models. Computer use is an application for them.
So the way they do computer use is to take the model, treat it as an API, and have the agent ask the model, “What do I do next?” Then the model produces the result.
That approach has three problems.
We were actually one of the first open-source agents using that paradigm, where the agent repeatedly asks the model and gets the result. That helped us become number one on the benchmark.
But we realized there are three problems with that framework.
First, it’s expensive.
For every move, every click, you need to take the whole screen and pass all of that screen information to the model.
Sometimes if you go to Wikipedia, it could be 100,000 tokens. Every step, you’re passing all of that information to a frontier model and asking, “What’s next? What should I do?”
It becomes very easy to spend $100 on a certain task.
It’s expensive.
Even though model prices are dropping in some areas, we can still see relatively steady pricing for frontier models because the model intelligence keeps getting better.
The second problem is speed.
For every step and every move, I need to package everything on the screen and pass it to the model. I may have to wait five or ten seconds for each move.
It’s slow.
The third problem is even worse: it’s not stable.
Foundation models are neural networks, so fundamentally they are probabilistic models.
Every time you ask the same question to ChatGPT, it can give you a different result.
Have you ever copied and pasted the same prompt into a large language model ten times? It can give you ten different results.
That reflects the fact that the model is probabilistic.
But in the agent space, we’re talking about an agent asking the model what to do at every single step.
Grace Shao (26:17)
And in your use case, you actually want exactly the same thing every time.
Versus wanting it to feel like a different human talking to me, your use case makes more sense if it gives you the exact same answer every time, right?
Ang Li (26:29)
Exactly. I just want the exact result.
I don’t want creativity in terms of which button to click.
If I deploy this agent on my server, I need to be able to anticipate what’s going to happen.
That isn’t necessarily the case with agents today.
For chatbots, that variability is okay because sometimes that’s what I want.
But these three problems are fundamental problems with using a single model to build agents.
We have a solution for these three problems, and it isn’t a single-model solution.
We have an architecture that we call a neuro-symbolic approach.
Thirty years ago, when nobody was talking about neural nets, everyone in AI was working on symbolic approaches: logic, reasoning, if-else statements, programs.
We try to combine the two.
We take some inspiration from what humans do.
For example, when I learn to ride a bike, the first time is really hard.
But if I ride a bike 100 times, it becomes muscle memory. I don’t even need to think about what I should do. I’m balancing myself without thinking.
Humans have this characteristic called the power law of practice.
As you practice, your efficiency at performing the job becomes extremely high, your consumption of brainpower becomes extremely low, and your reliability becomes extremely high.
Today’s agents powered by frontier models don’t have those characteristics.
Every time I ask them to do the same job, it costs me the same amount of tokens. I may pay the same $30 for one simple task. And every time, it can still give me some surprises and failures.
That’s the problem.
Our solution is basically something like note-taking.
Suppose I’m performing a job for the first time. At the same time, I write a playbook. That playbook is represented by code.
The agent has a notebook that remembers what it did to perform the job.
Next time, it refers to the notebook and tries to reproduce it.
Grace Shao (28:46)
That’s brilliant. Okay, so you save a lot of token usage in this process.
Ang Li (28:50)
Yeah.
One of our early experiments showed that for some tasks, we can save 90% of token usage.
So it’s 10x cheaper than a typical agent.
At the same time, you get higher reliability because most of the actions are code. It’s deterministic.
When I run it, I know the result.
It also helps make the system more transparent.
In an enterprise setting, when I look at the system, I know exactly what this thing is going to do every day.
I’m not worried about the AI suddenly jumping out of the sandbox and attacking other companies.
It’s much more controlled.
Those benefits help people adopt the technology in production environments.
Grace Shao (29:32)
That makes a lot of sense. That’s really interesting.
On the token-spend point, I want to ask you as a founder: how do you balance which models you use?
How do you decide which model to route to, and how do you decide on token spend?
Explain to us the pain point you face right now as a founder, and whether you have any solutions or suggestions for other founders.
Ang Li (29:55)
It really depends on the work you’re performing.
One observation is that there are basically two types of work.
One is content creation. The other is execution, which is what we focus on.
Content creation can be generating images, generating videos, generating code, generating apps.
I view programming and engineering as creation.
That kind of work requires the most intelligent models.
When a significantly better frontier model shows up, a lot of people will immediately move to it because they feel, “Okay, this model is smarter for writing my apps or doing programming work.”
It’s difficult to do the same type of job with the same efficiency using a much lower-tier model.
That’s one category.
The second category is something like, “I want to go on LinkedIn and see what’s going on, see how many people reached out to me, and click around.”
We don’t need an IMO gold-medalist-level model to do that kind of task.
For computer use, it’s highly likely that you only need a good open-source model to accomplish a lot of these jobs when it’s paired with the harness we’ve built, because most of the job gets translated into code.
When you don’t ask the model what to do at every single step, your dependency on the model becomes lower.
That means smaller models have the opportunity to produce equivalent performance for this type of work.
Computer use is a category that shows up a lot in non-technical corporate functions like finance, marketing, and GTM teams.
They go to different websites and extract information.
It’s not a small category. I would say 80% of what we do on a computer is actually looking at a screen and clicking things.
As a founder, you have to balance the tools.
Do you really need your GTM team to use the most expensive frontier model to click around LinkedIn or go to websites and extract information?
You don’t have to. You don’t need a PhD to do that kind of work.
It’s related to how you manage a team: put the right person in the right position, and put the right agent on the right work.
For computer use, we have an opportunity to drop costs dramatically. That’s technology you can already use to maximize productivity in non-technical domains.
For coding agents, I still think it’s important for engineers to use the best model because you’re doing content creation.
You want the agent to be smart in terms of interpreting your intention, getting feedback, and revising the code.
That’s an area where it’s currently very hard to say, “Let’s just use a worse model and the team will still have the best performance.”
My strategy is to make sure the team has all the tools available to them.
That space is moving very fast. Today we have one model; tomorrow we may have a better model.
As a startup, you want to move fast, and one way you move fast is by using the best technology available.
But I think the part most people overlook is that a majority of work doesn’t actually need the best model.
We have a solution for those workloads, and people should consider reducing their costs there.
Grace Shao (33:50)
That’s a very clear way of putting it, and I really appreciate that.
How do you view the way frontier labs are essentially incentivizing this race toward token maxing?
Over the last couple of months, we’ve also seen a lot of open-weight models push token costs down.
To your earlier point, that really changes the competitive landscape for agent companies like yours, which can be big consumers of tokens.
How do you view the incentives of open-weight models versus frontier models right now?
The frontier labs are clearly still chasing token maxing and more expensive tokens, while the open-weight models can continue compressing costs.
How do you see that dynamic playing out?
Ang Li (34:35)
Suppose you are Sam Altman or Dario. What would you do? Would you give your models away for free?
If I were managing a frontier lab, I would feel like I had no choice.
Their valuations are very high, and those valuations need revenue to justify them.
How do they get revenue?
People use their models, and they charge for token usage.
So you have two choices. One is to increase your price. The second is to increase usage.
That’s basically what’s happening.
People have been saying frontier-model prices will drop over time. But over the last few years, for the highest-end models, you still see premium pricing.
The most intelligent tokens can become more expensive.
Grace Shao (35:27)
They justify the premium.
Ang Li (35:29)
Exactly.
Then you have token maxing.
Those dynamics make sense in the current industry because these companies have extremely high valuations, they’re competing with each other, and they have to raise money and generate revenue to make the companies sustainable.
From our perspective, though, we’re not a frontier lab. We’re not a foundation-model company.
What we want is to serve normal people.
We want people to have more affordable and accessible options, so that everyone has the opportunity to automate their work.
We don’t want only rich people to have this opportunity.
We want to bring this kind of luxury to everyone.
Even for very tedious work, they should be able to automate it so people can be freed.
In that scenario, we don’t have to do token maxing.
We’re doing the complete opposite. We try to minimize the number of tokens you use.
And we don’t always have to use a premium model because this type of work doesn’t require a Math Olympiad gold-medalist model.
An open-source model can work.
That’s exactly why I find this category so interesting. It has the opportunity to allow everyone to offload their work.
But if we’re talking about coding, that’s a different story.
Coding agents need the best models. That means you have to pay the premium. They are going to be expensive.
That’s also why the frontier labs are still focused so much on coding agents. It makes sense for them because that’s one of the fastest ways to drive revenue.
I wouldn’t say it’s a simple question. It’s a complex societal problem where you have to consider economics and how startups work.
After all that analysis, their behavior makes sense.
Grace Shao (37:18)
That makes sense.
Touching on coding agents, last year was really interesting. It felt like the year of the coding-agent wars. Everyone was in an arms race, and that’s still ongoing.
Coding agents were arguably the first major breakout use case for agents.
And now we’re seeing more and more companies move into computer-use agents. Manus was one of the earlier examples.
Why do you think the market is mature enough now to actually consider computer-use agents as the next big area of competition and focus?
Ang Li (37:57)
Because there’s nothing else remaining.
I used to tell people: when the frontier labs start competing on computers, that means we’re getting close to AGI.
Grace Shao (38:09)
But how would you define AGI?
AGI feels vague, right? Everyone gives a different version.
If we’re really so close to AGI, it doesn’t feel like the world is suddenly going into doom, all our jobs are lost, we’re irrelevant, machines are taking over.
It doesn’t feel like that, which is the narrative Silicon Valley has been telling.
How do you define AGI?
Ang Li (38:32)
My definition is basically this: for my whole workday, I don’t need to sit in front of the computer. I just talk to my agents through my phone, and they do my job.
That’s it. It’s pretty simple.
I wouldn’t worry too much about humans losing their jobs.
You still need humans to sign off.
You still need someone to be responsible for the outcome. You can’t let the agent itself be responsible.
Grace Shao (38:56)
But then would your value still be as high?
If your job becomes so easy — and your job technically isn’t supposed to be easy as a founder or technical person — but it becomes much easier, are you still worth what you used to be worth?
Does your market value drop?
Ang Li (39:11)
You’ll focus much more on high-level strategy.
Even what kind of prompts you give the agents becomes important.
That’s one interesting thing about coding agents.
People say, “Okay, agents can write code.”
But if you tweak your prompt a little bit, the outcome can be very different.
If you understand the algorithm behind a coding agent, one of the first things it does when you send a prompt is extract keywords and search through your computer.
If you directly give it the right keywords, it becomes massively more efficient for the agent to do the work.
If you don’t give it the right keywords, or you use language that doesn’t appear anywhere in your files, it becomes less efficient.
So there are still subtle differences in what kind of prompts you give agents, what kind of decisions you make, and what kind of problems you choose to solve.
Those things become more important.
It’s kind of like the product manager’s job.
People keep saying product managers won’t be needed anymore.
But actually, that kind of strategy becomes extremely important once you remove all the other work from the process.
The good part is that this kind of work doesn’t require you to sit in an office anymore. It doesn’t require you to give up your physical freedom.
Grace Shao (40:29)
So should companies be training employees on how to prompt, essentially?
Is that going to become part of training — how to make you more efficient at your job?
Ang Li (40:39)
Yeah.
This kind of training is essentially what teachers have always tried to teach students in school: ask the right question.
I did my PhD at Maryland, and my advisor was a very senior person.
The biggest thing I learned from my PhD advisor was to ask the right question. Find the right problem to solve.
I spent five years constantly asking myself, “How do I find the right problem to solve?”
It’s an extremely hard problem.
I still remember him telling me, “If you find the right problem, 50% of the problem is already solved.”
Once you can write a problem statement clearly, 50% of the job is already done.
The remaining 50% is relatively easy because you follow that problem statement and search for results.
That’s happening in the AI space right now.
If humans have the capability to ask the right question, 50% is done. Agents can handle a lot of the remainder.
The reality is most of us aren’t trained this way.
Most education today trains people to solve problems.
Solve math problems. Take exams where somebody gives you a problem and asks, “Can you solve it?”
There’s no exam that asks, “Can you come up with an important problem and then solve it?”
I feel like the whole education system will evolve in that direction, pushing everyone to think, “What’s the right problem?”
I only have, I don’t know, 60 years of my life. What kind of problem is important enough for me to solve during my lifetime so I can make the maximum contribution to society?
Most people don’t have the opportunity to think about that because they are given homework every day at school, and then they’re given homework every day at work even after they become adults.
Now more and more people are starting to think: what should I prompt the agent to do?
Suppose I come up with an interesting prompt and it creates an interesting app. That’s the excitement people are getting from the current technology.
I view this as a positive change for society.
This could potentially take society’s creativity and innovation to the next level with these agent tools helping everyone.
But the important part is: how do we have a solution, an agent, that’s ready for everyone to use?
Not just people already in this camp. Not just people in Silicon Valley.
Can we let people running car dealerships, accountants in family businesses, dentists, solopreneurs, a five-person dental practice — can all these people use the technology and become maximally productive?
Then they can start asking bigger questions about their lives: what kind of thing can I do to make my life more meaningful to society?
Grace Shao (43:30)
That’s a very interesting way of putting it.
It makes me think that you’re essentially creating a tool that enables the average person to become a high-agency person.
Silicon Valley loves talking about high agency, but high agency, like you said, is also kind of a luxury.
Most people have agency. It’s just that you’re so bogged down by day-to-day mundane tasks, duties, and the work you need to do to earn wages that you can’t actually act on that agency.
If you have someone executing a lot of that tedious work for you, you suddenly have the mental capacity to put your energy elsewhere.
I really appreciate that.
I want to ask you another question.
Earlier this year, we saw the “SaaSpocalypse,” and that was quite wild. It was a wild ride for the market.
More recently, people have started saying that reaction wasn’t very sensible.
SaaS companies have existed for decades because there’s industry know-how, workflows, processes, and systems in place. These things can’t necessarily be replaced overnight by something vibe-coded.
However, what you’re saying is that the agents you’re building can actually operate these SaaS products on the desktop.
How do you view SaaS companies going forward?
Over the last six to eight months, we’ve definitely seen a bit of a flip-flop.
Now people are saying, “Actually, ServiceNow cannot be easily disrupted. My God, how could the market have reacted that way?”
What’s your sensible take here?
Ang Li (45:09)
I feel like the industry has a pattern of viewing something as just one thing.
I always say that when you look at something, there are usually two angles.
When we look at SaaS, we also have two categories.
The first category of SaaS companies stores the data. They are the systems of record.
They record things in their own databases and serve that information to users.
Another set of SaaS companies doesn’t really store the core data. They build a portal on top of other systems.
So there are two types of SaaS companies.
My principle is that infrastructure tends to stay.
Over the past 50 years, infrastructure doesn’t just disappear. Humans are really good at building layer upon layer of infrastructure.
The first set of companies, the ones that have the data, are infrastructure.
Data is kind of like electricity in the digital world.
The data is important because it powers further innovation.
For example, Salesforce has the data.
That’s important.
Those companies won’t simply disappear because they form part of the infrastructure.
Agents can sit on top of these systems of record.
Agents can replace a lot of the human-made portals in between because now everyone can create a portal with an agent, in real time.
You don’t necessarily need a company to spend 10 years building a portal and then sell it to customers.
You can have an agent build a portal on top of the system of record in a day.
So that second type of SaaS company may face existential risk because it’s competing with the agent layer.
At least half of them are fine. There’s nothing to worry about.
This also goes back to the question people always ask: will GUIs still be there? Why not rebuild everything around APIs?
We should view it partly as an infrastructure problem.
If software and its GUI have existed for 20 or 40 years, we should probably view that as infrastructure.
Then you can build agents on top of it to modernize that infrastructure.
But if you have a SaaS startup that built a portal used by a small fraction of users, and the portal has only existed for three years, it may not be strong enough to become infrastructure.
If it isn’t strong enough to be part of the infrastructure, then it becomes easier for a company like ServiceNow or Salesforce to say, “Okay, I’ll build an API for that.”
Then the agent can connect directly to the underlying system, and your portal may no longer be useful.
We really have to identify what is truly infrastructure and what isn’t.
Even within GUIs, there are two parts.
Long-lasting legacy software that is difficult to move away from has already become infrastructure.
But modern software produced by many Silicon Valley startups hasn’t necessarily achieved that status yet.
That kind of software can potentially be removed by a system-of-record company producing an API and connecting directly with agents.
Grace Shao (48:46)
I see. That makes sense.
I just want to ask you two questions I always ask everyone.
One is: what do you think people really misunderstand about your industry? In this case, CUAs.
And the other one I’m going to throw to you now so you can think about it: what is the differentiated view you hold?
It could even be something non-AI-related, if you feel really strongly that the Earth is flat or something.
Ang Li (49:08)
I think on the first question, I’m always being misunderstood because people keep telling me the future will be all APIs.
I hope the future is all APIs, but I think it’s actually impossible for society to become entirely API-based because APIs are not transparent.
It’s hard for you to see what’s going on.
GUIs are more transparent.
You can see exactly what the agent is doing.
I actually feel safer when I look at a screen and I can see, “Okay, the agent is moving the mouse, clicking on a button.”
I appreciate having the opportunity to go in and stop it if I see something going wrong.
I also appreciate the fact that it isn’t necessarily too fast.
If the entire digital world becomes 100x faster than human processing speed, I don’t think that’s necessarily a good thing for us.
It’s nice that everything becomes fast, but you don’t want humans to become overwhelmed by intelligent systems.
That’s one of the things frontier labs are dealing with.
Humans have a limitation on processing speed, so we want to maintain some kind of balance.
That’s why I feel like computer use is actually a relatively safe road for society to move down.
It operates around human speed. Sometimes it’s even slower than a human.
Grace Shao (50:35)
That’s very interesting.
Ang Li (50:37)
Yeah.
And we won’t be spamming the whole internet because it’s slower.
Worst case, I’m just performing a task like a human. Why is that necessarily a bad thing?
People worry about spam, fraud, and security because you could have something processing at 100 times the speed of a normal human.
That can overwhelm the entire infrastructure of the digital world.
That’s why I feel like working on computers isn’t just a technology problem.
It’s also about asking: what’s the best way to make sure AI is safe, manageable, and transparent while still solving real problems and freeing people’s time?
These views didn’t all come from day one.
Over the course of developing the technology and talking to customers, we gradually realized that while frontier labs worry about their agents being unsafe, we’re actually pretty confident deploying this into the real world.
We think it can benefit people without causing damage because people can always look at what it’s doing, and they have time to react if something goes wrong.
Grace Shao (51:50)
That’s really interesting.
Your thoughts are probably also evolving as these scenarios play out in real life.
It’s interesting that you’re saying even what looks like a disadvantage — the latency, the lag, the controlled environment — can actually become an advantage in this scenario.
Ang Li (52:08)
An advantage, yes.
The frontier labs are now saying, “Okay, maybe we should pause. We should slow things down.”
For us, we’re basically maintaining this speed already, and it isn’t creating those kinds of problems.
Grace Shao (52:24)
That makes a lot of sense.
All right, last question. What is one differentiated view you hold? Something wild or non-consensus?
Ang Li (52:32)
I’m not sure if this is really a differentiated view.
My view is that in AI development there are two camps.
A lot of people are focused on the first camp. We’re in the second.
Everything is ultimately about how to do human work.
The first camp is trying to do more and more high-end work, gradually moving towards scientists’ jobs.
Initially, everyone thought AI was going to replace all human work, and a lot of blue-collar workers would lose their job security.
But the reality is that some of the easiest jobs to replace may actually be scientific jobs. Even AI researchers’ jobs.
That’s an interesting pattern.
When I talk to my friends — we used to be researchers training models — I ask researchers at big labs, “Do you think going into a portal and clicking around is easier to automate, or replacing your model-training pipeline is easier?”
Most people will say replacing the model-training pipeline is much easier because the whole procedure is so routine.
There’s already a playbook.
Suppose you want to train a model. There’s already a playbook. You follow the playbook and do it every day.
It’s actually a very tedious job.
I used to be on the hiring committee at DeepMind, and we would hire machine-learning engineers into the company.
Everyone was so excited. They thought, “DeepMind is exciting. It’s a frontier AI lab. If I go there, I must be doing something incredibly exciting.”
That was before they joined.
After they joined, a lot of people realized, “Why am I always cleaning dirty data? Why am I always doing tedious engineering work to build around the system, maintain data quality, and feed it into the model? Am I supposed to be training models? Am I supposed to be changing the model architecture?”
But in reality, for many AI researchers, you spend less than 5% of your time actually looking at the model.
The majority of your time is spent on tedious things like data cleaning, data labeling, organizing data, and moving files around.
That’s the interesting part.
When I started working on computer use, initially I saw it as a challenging problem that nobody else was working on.
Over time, I realized it might actually be one of the hardest problems, even compared with these so-called high-end jobs.
If you look at the frontier labs’ strategies, they are moving toward high-end work: having AI write code, having AI train models.
Basically, they’re trying to automate more and more of an AI researcher’s job.
We’re trying to come from the other direction.
We’re looking at all of this repetitive work and asking: if it’s so repetitive, why is it actually harder to automate than an AI researcher’s job?
The reason is that most of the AI researcher’s work is already handled through code.
When you handle something through code, the data is structured.
But when you look at a screen, the screen is extremely unstructured.
It’s similar to understanding video content. You have a real human in the real world, and understanding that kind of environment is a much harder problem than solving a coding problem.
Grace Shao (56:03)
That’s really, really insightful. Thank you so much for your time.
I just want to end on one thing you’re saying, which is basically that nobody’s job is as glamorous as it looks from the outside.
It reminded me of when I still worked in broadcast TV.
People would say, “You must just be talking to four or five Fortune 500 CEOs and looking pretty on TV.”
No.
Most of the time, we were squatting in a corner waiting for the guest to come out, eating takeout — or not getting to eat or go to the bathroom for 10 hours — and just waiting around.
We were waking up at three in the morning, doing research and makeup at the same time.
Nothing is as glamorous as it looks from the outside.
So much of the work is actually preparation.
Ang Li (56:39)
Exactly. Yeah, exactly.
It’s like founders.
Grace Shao (56:43)
Thank you so much for your time today.
Ang Li (56:46)
Thank you so much. Nice chatting with you.
AI Proem is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.
Get full access to AI Proem at aiproem.substack.com/subscribe
Fler avsnitt
Visa alla avsnitt av AI Proem PodcastAI Proem Podcast med Grace Shao finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.