Overview
◆ lecture.logStanford CS146S2026-10-08
CS146S · Lecture

How always-on
agents work

What makes an always-on agent tick

~14 min read

Speaker
Lee Robinson
Course
Stanford CS146S
Date
2026-10-08
Runtime
45:17

Speaker now at SpaceX (ML, per his website; formerly Cursor). Product: Bot; GrokBot in the talk. No Q&A recorded.

Thesis

An always-on agent is the basic agent loop moved onto a server. The loop: “model + tools in a loop”: the model calls tools until none are needed; CS146S students built it in week one. On a server:

  • events wake it; it runs a turn, then stops;
  • a cloud computer that sleeps;
  • all state in a database.

What makes it work: better models, and the infrastructure and context engineering around the loop.

◎ panorama5 acts · 13 ch00:00–45:17
Question

How does a loop that stops when the terminal closes become a colleague that works for months?

  1. I00:00–10:04
    StartWhy

    As models improve, the agent goes from "waits for approval" to "works while you steer"; the interface gets simpler.

  2. II10:04–18:07
    BackboneAlways on

    Nothing runs when idle: events wake it, a turn runs and stops, state is saved, crashes resume from the last step.

  3. III18:07–27:24
    HarnessHow it works

    One tool talks to you; other tools load on demand; helpers take heavy work; four safety defenses.

  4. IV27:24–34:27
    ContextStay sharp, stay cheap

    Stable prefix, compaction while the cache is warm, layered memory: long runs stay sharp and cheap.

  5. V34:27–45:17
    FutureWhere it goes

    Every failure becomes an eval; build as little as possible; learn the new building blocks and fundamentals.

CH 0000:00–02:34S01–S05

Your agent vs. an always-on agent

From the week-one loop to always-on

Same loop, new way to run it

The week-one agent (S03: "You built this in week one") stops when the terminal closes. This lecture: keeping the same loop running.

Key points

The week-one loopS03

messages = [system_prompt, user_message]
while True:
    reply = llm(messages, tools=TOOLS)
    messages.append(reply)
    if not reply.tool_calls:
        print(reply.text)
        break
    for call in reply.tool_calls:
        messages.append(run_tool(call))
Visuals
ComparisonS04
AspectWEEK-ONEbasic loop in the terminalALWAYS-ONagent on a server
RunsIn your terminalOn a server: survives crashes, deploys, and a closed laptop
ComputerYour laptopIts own cloud computer with a desktop, browser, and files
MemoryOnly the messages listSaves state, memory, and skills; summarizes old messages
TalkingPrints every replyMessages you only via a send tool, so it can stay quiet
StartsWhen you typeOn events: messages, schedules, Slack, phone calls, finished work
LifespanOne task, then exitsWorks for months with the same memory
From the talk
  • Like a colleague: knows when to stay quiet, works in the background, returns only with important updates. Goal: weeks or months of work.
slides · S01–S05 · 5
CH 0102:34–06:13S06–S08

How developers work with AI

How we got here

From approving to steering

AI does more; humans shift from approving each step to steering anytime. The interface: as simple as texting.

Visuals
Four erasS06

2022–23Copy and pasteChat in a browser tab; move code by hand.

2024–25Agents in the terminalIt edits files and runs commands; you approve each step.

2025–26Apps built for agentsMany agents run in parallel, in the background or cloud; you review results.

Nowalways-onYou text an agent with its own computer, tools, and memory; it keeps working.

The higher the step, the more the AI does.

Before / Now tracksS07
BEFORE
Approve?Approve?Approve?

Before: every step stops at "Approve?" and waits.

NOW
SteerSteerSend inviteauto-review

Now: interjections ("Also check the Slack thread", "Actually, make it Thursday") join one workflow; auto-review approves the calendar invite◐.

Key points
  • Their implementation: Bot (S08). Bots left (Sales Outbound, Inbox Manager...), chat right; lacking a Salesforce login, it asks you to sign in on its computer.
From the talk
  • Many companies call "a separate model reviews every command" auto review. The speaker: safer than a tired human approving hundreds; a blind test would likely favor the model.
slides · S06–S08 · 3
CH 0206:13–10:04S09–S11

Six new model capabilities

What changed in the models

Knowledge lives in files

Six new model capabilities enable always-on. Knowledge sits in plain files; as the slide says, "Everything is computer".

Key points

Six capabilitiesS10

  1. Work for hours, following instructions throughout.
  2. Hand off to helpers, each with a fresh context.
  3. Use real tools: shell, files, browser, MCP.
  4. Use a computer, even apps with no API.
  5. Keep knowledge in files.
  6. Learn by watching: show it once, it saves a skill.
Visuals
File treeS11 · simplified example
memory/profile.mdAlways in the prompt.
memory/log/2026-10.mdDated facts, one per line.
skills/*/SKILL.mdLoaded only when needed ("Use this when...").
routines/morning-briefA prompt that runs weekdays at 8:00.

Why files

  • Models are extremely good at the shell.
  • Plain text both you and the agent can read and edit.
  • State a preference once; it goes into memory.
  • Skills load on demand, so many are cheap.
slides · S09–S11 · 3
CH 0310:04–15:25S12–S18

Inside an always-on agent

Lifecycle and the three-part system

When nothing is happening, nothing runs

Always-on doesn't mean always running: when nothing happens, nothing runs; everything is saved, so any trigger brings it back.

Visuals
Lifecycle simulatorS13 · S15
LANENOW
ProcessRuns when triggered; stops after 2 idle minutes
○ stopped
ComputerRestored when a tool needs it; snapshots and stops after a few idle minutes
○ asleep
StateDatabase and storage (conversations, memory, routines, skills), always kept
● always on
  • The computer never sleeps mid-turn.
  • If a routine doesn't need the computer, it stays asleep.
  • No tool calls while the process isn't running.
  1. t0—
Wake-up routerS14
RouterDrops duplicates, picks the conversation
Conversation queueOne turn at a time
Anything worth saying?
Yes: send to the right place
No: stay quiet

Every wake-up records where it came from◐ —

Three-part systemS16
Apps & channels
Desktop
Mobile and web
Slack
Email / webhooks / schedules
Phone calls and meetings
Server
Entry point
One queue per conversation
agent loop
Model
Database: load and save
Cloud computer
File and connector server
Command runner
A screen and Chrome per agent◐

Conclusion: every app shows the same conversation.

Key points
  • Cloud computer (S15): one Firecracker VM per user; Debian container kept alive by a supervisor. Files, installed tools, browser logins persist across sessions.
  • Why the loop runs on the server (S17): laptop closed, work continues; seamless computer–phone switching; no App Store review cycles; months-long agents can't live on devices◐.
  • Key point (S18): state lives on the server, so the VM can be rebuilt anytime (e.g. for a Linux 0-day) without stopping the loop.
From the talk
  • Everyone's computer always on would cost too much; "there's not enough computers in the world". At SpaceX, people often send bots to take meeting notes.
slides · S12–S18 · 7
CH 0415:25–18:07S19–S21

Messages, crashes and the queue

Reliability

Messages, crashes, and cutting the line

Reliability rests on three things: unique IDs and sequence numbers on messages, a durable workflow per turn, a priority queue for new messages.

Key points

What happens when you hit sendS19

  1. Sending: shown locally at once with a unique ID; the server drops repeated IDs, so retries are safe.
  2. The turn: joins the queue; one turn at a time per conversation; every step saved to the database.
  3. Watching: the server sends only a "something changed" signal; apps fetch the data, so a lost signal is harmless.
  4. Ordering: every update has a sequence number (41, 42, 43); gaps reload from the database◐.
Visuals
Crash simulatorS20
Job queuestarts over at Step 1
durable workflowpicks up at Step 3

Repeated steps:0 / 0

Note: steps are recorded and idempotent (running twice is harmless); open-source engines like Temporal include timers and cancellation◐.

Cut the line and mergeS21

In a 1:1 chat, "Make it Thursday" or Stop skips the line and interrupts the running turn.

Group chats, Slack, routines, helpers: back of the line; 3 quick messages merge into 1 turn.

slides · S19–S21 · 3
CH 0518:07–19:59S22–S24

One tool for talking to you

Harness ①: a single conversation tool

SendToUser

The agent talks to you through one tool, SendToUser; anything needing your OK gets a purpose-built UI control.

Visuals
HubS23
12 tools for doing work
ShellFilesBrowserDesktopMemorySkillsHelper agentsCloud agentsApp connectorsWeb searchImage generationScreen recording
SendToUser
DesktopWebMobileSlackEmailVoice
  • Thinking stays private.
  • Nothing new, no message.
  • Rich messages: files, multiple choice, secure password entry.
  • The last message ends the turn.
UI controlsS24

Sign-in

••••

1Password fills the form; the password never reaches the model.

Payment

Waits until you tap Allow or Deny.

Email draft

You choose Send or Discard.

Reaction◐

From the talk
  • The speaker: the biggest difference from the week-one harness (maybe also similar products) is this single tool between client and server.
slides · S22–S24 · 3
CH 0619:59–26:27S25–S29

Tools and helpers

Harness ②: tools and helpers

Try the cheapest option first

Tools start as names; information sources escalate by cost; helpers do heavy work and report back briefly.

Visuals
Context usageS25 · illustrative
Every tool loaded up front

Tool details 62 · free 8

Names first, details on demand

Tools 8 (details 6, names 2) · free 62

System promptTool detailsTool namesConversationFree space

1 cell = 1%, per the slide's illustrative proportions. Also called dynamic tool discovery / tool search tool.

Cost ladderS26
What it already knows
An app's API (connector)
web search
Its own signed-in browser
The full desktop
Ask you

Slower, more expensive◐

Rule: connector broken → don't work around it with the browser; tell you.

Main agent + helpersS28
Main agent
Big task
Four kinds of helper:General (shell, web, connectors)Computer useVideoCloud coding◐
Short report

The main agent answers quick questions, hands off big tasks, gets a short report back.

Key points

Three ways to use a computerS27

  • Read the page structure: list buttons and links; fast and reliable.
  • Click on pixels: a 1280×800 screen, for desktop apps and system dialogs.
  • Ask a person: hand over at logins, two-factor codes, captchas, payments.
  • A helper does the clicking, keeping screenshots out of the main context.

4 tips for tool designS29

  1. Don't add or remove tools between turns.
  2. Keep each tool's description and schema under 4,000 characters.
  3. Save state with a dedicated tool, so it can be reviewed.
  4. Errors say what to do next: "The computer is asleep. Use Shell or Read to wake it, then try again."
From the talk
  • Each helper is a durable workflow that survives crashes, so the product "feels like it has infinite context".
slides · S25–S29 · 5
CH 0726:27–27:24S30

How we keep it safe

Safety

Four defenses

Four defenses: model review, source labels, ask before sending, secrets kept apart.

Key points
DefenseRule
Automatic reviewA separate model checks risky actions first; blocked ones go to you for approval; no answer in 15 seconds keeps them blocked
Track the sourceEmail and webhook text is labeled untrusted; memories are information, not instructions
Asks firstNo messages to others unless you asked; new routines need your OK
Keep secrets hiddenPasswords go into environment variables, never the model; actions on your laptop need approval
Visuals
Auto-review flowS30
Risky action
Model review
PassRun
BlockedAsks you15 sec countdownNo answer: stays blocked
slides · S30 · 1
CH 0827:24–31:31S31–S34

Keep the start of the prompt the same

Context engineering ①: prompt caching

A cache miss is a bug

Most tokens of a long-running agent should hit the cache: fixed prompt start, append only at the end, every cache miss treated as a bug.

Visuals
Cache stackS32
Stays the same
Tool definitionsSystem prompt (with memory and routines snapshot)First message (environment, tool names)
History
Only appended at the end
Newest message
Time, change notes, your message

Cache hitcache miss

Key points
  • Main context keeps only pointers or summaries (S33): long outputs to files, screenshots and heavy work to helpers, old messages to summaries, instructions to skills.
  • Treat a cache miss like a bug (S34): prefix unchanged until the next summary; deploys send running conversations a note, not new tool text; fixed order; per-section hashes; cache warmed as you type.
From the talk
  • OpenClaw's January boom brought many cache misses and many fixes. A/B tests must not reorder the prompt either.
slides · S31–S34 · 4
CH 0931:31–34:27S35–S36

Summarize while you're away

Context engineering ②: compaction and memory

Summarize while the cache is warm

Summarize while the cache is warm. Four memory layers: Profile always in the prompt; Log adds only the 30 most relevant entries.

Visuals
Compaction timelineS35
01A big turn ends (over 100k tokens)
02About 10 minutes pass (cache still warm)
03Write a summary (reuses the cache)
04Your next message starts from the summary
  • Skipped if background work will wake the agent soon.
  • Last few messages kept word for word.
  • Overlong turns can be summarized mid-turn.
Four memory layersS36
ProfileCore facts, always in the prompt: up to 100 facts, 4,000 characters.
LogDated events; only the 30 most relevant recent ones.
NotesShort-term details that fade quickly.
The restFound with a search tool.
  • Important facts rank higher; older ones slowly sink.
  • It can only remove facts it was shown.
  • Facts about you are shared by all your agents.
  • Memories are information, not instructions.
From the talk
  • All compression is lossy; there are probably "entire PhDs just for this problem".
slides · S35–S36 · 2
CH 1034:27–37:43S37–S38

Product and model improve together

Product and model evolve together

Every failure is an eval

Models are trained for the product. Each failure becomes an eval: the inner loop fixes the harness, the outer loop trains the next model.

Key points
  • Three training stages (S37):
    • Pre-training: language, code, general knowledge.
    • SFT: learns the harness: tools and formats, when to reply quickly.
    • RL: practices long instructions, skills and memories, clicking the right spot (grounding), working with other bots.
  • Inference is tuned separately: the main agent replies in seconds; helpers reliably run for hours◐.
Visuals
Two loopsS38
Inner loop (every day)
  1. 01A bot gets something wrong
  2. 02Turn it into an eval
  3. 03Fix the harness, prompt, or inference
  4. 04Check the evals
  5. 05Ship the fix
Evals guide training
Outer loop (every few months)
  1. 01Train the next model on the harness and evals
  2. 02Swap it into the product
  3. 03People use it in new ways (e.g. group chats full of bots)
  4. 04New behaviors to evaluate and train for
slides · S37–S38 · 2
CH 1137:43–43:22S39–S42

Where this is going

Where it goes from here

Delete the product

Delete the product: assume models get much better, build as little as you can around them, remove what they no longer need.

Key points
6–12 month predictions (S40)Matching open problems (S41)
Longer projects (months)How to test agents that run for weeks
Better memoryCorrect memory: update, forget, never make things up
More powerful toolsPrompt injection: text in email and web pages written to trick it
Simpler interfaces (maybe voice only)When to speak: too much is noise, too little hides progress
Skills learned on the jobCoordination: agents sharing a computer and memory = distributed-systems problems
Knowledge written downCost: idle nearly free, active cheap

Rows pair up by theme.

Visuals
Building blocksS42
Every web app had
A serverA databaseA cacheA CRUD interface
Every agent product will have
durable workflows04 Can use a computer0306 Memory and skills0209 Triggers and scheduled jobs03 Connectors and auth0506 Text and voice channels05 Payments and identity0507 Evals and observability10

Chips = related chapters

From the talk
  • "The models today are the worst they'll ever be": the UI may need rebuilding within six months. Spoken addition: security against bad actors taking over agents.
slides · S39–S42 · 4
CH 1243:22–45:17S43–S44

What to take away

Takeaways

Architecture compounds

As code gets easier to write, knowing how the pieces fit together matters more.

Key points
  • Architecture compounds: when code is cheap, the right architecture keeps you faster than everyone else.
  • Build your own: start from the week-one agent; add a server, a computer, memory. Open-source always-on agents are great to learn from.
  • Learn LLM fundamentals: like knowing databases when building a website◐.
Visuals
"Build your own" checklistSuggested order0 / 6
From the talk
  • He still reads his code, especially to understand the architecture he ships.
slides · S43–S44 · 2
SUMMARY

Summary

Summary

6 principles across the lectureSynthesis

PrincipleChapters
State is the only truth; everything else is disposable03Lifecycle04Reliability
Silent by default; speak only with reason03Lifecycle05SendToUser
Main context is the scarcest resource06Tools & helpers08Caching09Memory
Append, don't edit04Reliability08Caching
Try the cheapest, most controllable path first06Tools & helpers07Safety
Every failure is an eval10Flywheel

The speaker's close: assume models get much better; build as little as you can around them. As code gets easier, how the pieces fit matters more.

GLOSSARY

Glossary

Glossary

glossary25 terms
TermMeaningExample
Always-on agentAn agent whose state lives on a server, woken by events, able to work for a long time; nothing runs when idleBot / GrokBot
HarnessThe runtime around the model: tools, prompt, loop, state managementGrokBot harness
Agent loopThe loop where the model calls tools, gets results, and calls againS03
RoutineA prompt run on a schedule or by a webhook, like a cron jobWeekday 8:00 morning brief
RouterComponent that drops duplicates and routes events to the right conversationS14
TurnOne run of the agent over a batch of messagesS19
FirecrackerLightweight VM technology, easy to snapshot and restartThe agent's cloud computer
Thin client, thick serverArchitecture with a light client and a heavy serverS17
Durable workflowExecution where every step is recorded and a crash resumes from the last stepTemporal
IdempotentRunning a step twice gives the same result and does no harmS20
SendToUserThe agent's only tool for talking to youS23
MCP / ConnectorProtocol for models to call external apps / a tool that connects to an app's APIGmail, Plaid
Dynamic tool discoveryGive tool names first; load details when usedS25
Computer useThe model looks at the screen and uses mouse and keyboard1280×800 screen
Accessibility treeStructured list of page elements; uses fewer tokens than screenshotsS27
GroundingMapping an instruction to the right spot on screen"Clicking the right spot"
Subagent / helperA sub-agent with its own context, sent out by the main agentVideo helper
Auto-reviewA separate model reviews risky actions before they run; if blocked, it asks youNo answer in 15 s keeps it blocked
Prompt injectionInstructions hidden in outside text to trick the agentPhishing email
Context windowMaximum tokens a model can handle at onceE.g. 200k tokens
Prompt caching / cache missReusing an identical prefix to cut cost / the prefix changed, so full priceS32, S34
CompactionCompressing a long conversation into a summary; lossyS35
EvalTest case that measures model or product behaviorS38
SFT / RLSupervised fine-tuning / reinforcement learningS37
Delete the productAssume models get better; build as little as possible and delete what is no longer neededS39