Working on agents
What I built at TikTok since early 2025, next to what the industry was doing.
2025
In the industryEarly 2025
- I paid for Devin, and it was the first time I really used an agent. AI was no longer just models. Of OpenAI's 5 levels of AI, the first 2, chatbots and reasoners, were work for the people who train models. Level 3, agents, was arriving, and as a software engineer I could build it.
- So I left TikTok's campaign work and conventional software engineering behind and went all in on agents. I started TikTok's internal general-purpose agent, first on my own with LangChain and the newly announced MCP.
- The ReAct paper led me to CodeAct, the research behind OpenHands: an agent acts by writing and running code. I rebuilt the demo on OpenHands. It planned, browsed, ran code, and searched the web, and showed each step as it ran.
- Devin is the showpiece for autonomous agents: it finishes the job instead of answering.
- Andrej Karpathy names “vibe coding”, and Claude 3.7 Sonnet ships with a preview of Claude Code.
- Manus launches and everyone rushes to study it. OpenManus, an open-source take, appears within days.
Apr – Jun
- Connected the agent to TikTok's operations data and actions through MCP.
- Pushed the OpenHands version further: a new UI, browsing, file upload, and MCP tools the agent could write for itself.
Jul
- After Manus, open-source general agents like Suna set the pattern. We moved our agent onto Suna's framework as the base for production. I added Kimi K2 3 weeks after its release, and we ran self-hosted open models alongside hosted ones.
- Surveyed the agent infrastructure stack: sandboxes (E2B, Daytona, Modal), browsers built for agents (Browserbase, Browser Use), and model routing (OpenRouter).
- Started a knowledge base, inspired by NotebookLM. Weighed retrieval approaches from plain RAG and services like Vectorize to Microsoft's GraphRAG, and explored long-term memory with mem0.
- Worked on prompt optimization with tools like Dify, and traced agent runs end to end with Langfuse.
- Studied how agents are evaluated: SWE-bench and Multi-SWE-bench, GAIA, τ-bench, EvalPlus, Chatbot Arena, and Video-MME for video.
- Moonshot open-sources Kimi K2, a model built for agentic tool use.
- ChatGPT agent combines research and action. Dia opens a skill gallery for reusable browser workflows.
Aug
- Added data tools for creators (profiles and side-by-side comparisons) and slide generation.
- Evaluated the ways an agent can get information from the web: search APIs (Exa, Tavily, Brave) and crawlers and readers (Firecrawl, Jina Reader). Built web and image search on Brave, plus YouTube search.
- Started moving from one agent that does everything toward capabilities people can pick.
- GPT-5 ships.
Sep
- Shipped skills in our agent, inspired by Dia: reusable workflows that people create, test, version, publish, and share.
- There is still no shared format for agent skills.
Oct
- Rendered interactive widgets within days of the Apps SDK. I built the widget host, widget state that survives a reload, and fullscreen widgets.
- Added video analysis for a single video or a whole batch.
- OpenAI launches apps in ChatGPT and the Apps SDK.
- Cursor ships Plan Mode. Days later, Claude Code adds an interactive question tool, so the agent asks before it guesses.
- Anthropic introduces Agent Skills and the SKILL.md format.
Nov – Dec
- Built plan mode for our agent, inspired by Cursor and Claude Code. The agent proposes a plan, and the user reviews it before anything runs.
- Moved the chat to a new thread interface and built a set of TikTok widgets: video lists, comment analysis with CSV export, creator profiles, search, and command output.
- Gemini 3 leans into generative UI. Cursor 2.1 lets the agent ask clarifying questions through an interactive UI.
- Agent Skills become an open standard.
2026
In the industryJan – Feb
- Moved our skills to SKILL.md.
- Made runs pause for a person: when a tool needs the user to sign in or approve, the run stops and resumes where it left off. Durability became a design constraint for our cloud agents.
- My first merge from a coding-agent branch.
- OpenClaw goes viral, and local, always-on agents become the pattern to copy.
Mar – Apr
- Started a desktop agent app, choosing pi as its agent core. Within weeks the entry point looked settled: Codex kept getting stronger, and Doubao Work was taking shape inside the company. We stopped the app at the end of April and stopped building a general agent.
- After OpenClaw took off, CLIs became how agents reach tools. I contributed preview-deployment commands to TikTok's internal CLI, packaged Lark (TikTok's Slack), operations data, and deployment as CLIs and installable skills, and wrote a skill that teaches agents to design agent-first CLIs.
- Built a deploy CLI that coding agents install and use themselves: generated code becomes a shareable URL in one command. This was the start of Sites.
- Codex, Claude, and Doubao Work take shape as general agents, and people start their work there.
May – Sep
- Sites grew from that CLI into its own product, TikTok's Sites, a month before ChatGPT Sites launched. Any team's coding agent builds a Site already wired into employee sign-in, internal data, Lark, and AI models, from a quick demo page to a production back office. The deploy CLI reached 1.0 and became part of Sites.
- Every Site became agent-native. Each can bind its own agent and exposes its user actions over MCP and WebMCP, so every Site can also be a ChatGPT plugin. New Sites get a database automatically, plus file storage and scheduled tasks.
- Agents can publish a Site, read its logs, query its data, and roll it back through the same APIs people use.
- Anthropic's Thariq Shihipar argues that HTML beats Markdown for agent output, and the idea spreads fast.
- OpenAI launches ChatGPT Sites in public beta: ChatGPT builds, publishes, and hosts websites and lightweight apps.
- WebMCP lets agents in the browser call a page's actions directly.
- OpenAI DevDay's program includes a session on Sites and the next generation of web apps.