Skip to main content

Chat demo walkthrough

This is a script, not a reference. It takes a seeded install from “the chat page exists” to “I have seen the conversations column, both kinds of chip, a questionnaire, an attachment, ⌘K and the CLI half — and I have seen four of them refuse”. Each step leaves behind what the next one needs, so run it in order the first time. A demo that only shows the happy path teaches nothing about the product, so the refusals are a step rather than a footnote. The reference pages behind each step are Chat & Sessions, Ask forms, Conversation Search and Chat Surface Limits. This page assumes none of them.

The running order

Before you start

State these to your audience. Three of them are the reason a demo falls over. A seeded install, and a login. crewship seed prints the account it creates. It makes seven demo agents across four crews — comfortably under the twelve-agent fan-out cap, so nothing in this walkthrough disappears from the conversations column. On a workspace with more than twelve agents, read the thirteenth agent first: the existing conversations of agents past the cap never appear in the column at all, with no row standing in for them, and this walkthrough would be demonstrating that instead. The CLI, pointed at the same server. Nothing defaults to it:
A model credential, but only for two moments. Steps 1 and 5 want the agent to actually reply — step 1 so the thread has a second turn, step 5 so there is something to find. Everything else here runs without a model: chips, the form sheet, the rendered preview, the upload, the provenance line and every CLI command are client-and-server work. An agent with a crew, and a writable output tree. An attachment lands in /output/<agent-slug>/attachments/…, so the agent must have a crew_id — an agent with none answers 404 Agent not found on upload, not a friendlier error. It does not need a running crew. The server writes the bytes host-side first, and only replays the write through the crew container when the host write comes back with a permission error — which is what happens once a crew has been provisioned and chowned its trees to uid 1001. So a never-provisioned crew uploads fine, and a provisioned-then-stopped crew answers 409 with the remedy in the message. Know which of the two you are standing on before you promise either. A form configured on the agent. The questionnaire step has nothing to show until one exists. Step 0 checks.

Step 0 — Check what the seed gave you

1

Read the agent

You are looking for two things near the bottom:
The demo seed grew these in parallel with this page. If crewship agent get casey shows neither, your install predates that work — the rest of the walkthrough still runs, but configure them yourself first with the two commands below. Everything after step 0 is written against the seeded casey / bug-report pair.
2

If they are absent, set them yourself

This is also the shortest demo of how they are configured. Both columns ride the ordinary agent PATCH — there is no prompts endpoint and no forms endpoint.
Suggested prompts are one question per line, at most 8 lines of at most 120 characters. Ask forms are a JSON array — at most 4 forms of at most 6 fields — and in practice always @a-file.json, because a form with five fields and a template is not something anyone types on a command line.
3

Note who did not get any

riley deliberately has neither. Keep the tab open — the contrast is step 2.

Step 1 — /chat is a surface, not a drawer

Open /chat. It is not an index. It picks up where you left off: the freshest conversation in the workspace opens immediately, the address bar is rewritten to that conversation’s /chat/<agent>?session=<id> shape, and a live WebSocket is open from the moment the page settles. On a workspace with no conversations at all you get “Pick a conversation” instead.
1

The left column

The column lists conversations, not agents. The agent is an attribute of a row — its avatar, and its name on the row’s second line — the way a sender is an attribute of an email, not a folder you have to open first.Top to bottom it is the same chrome as /issues and /routines, and deliberately: a Search + Filter toolbar, one bucket section — here Show, with Direct, Routines and Issues, each carrying its count — then the list. Conversations, newest activity first, grouped by day once there are enough of them to need headings. The + beside that heading starts one.
2

The scopes

Click Routines. The list you were reading is replaced by routine steps, stacked under the routine that ran them.This is the one control on the column that is not a filter. It is a fetch parameter?kind= on GET /api/v1/agents/{id}/chats, which narrows inside the statement, before its LIMIT. That distinction is the whole reason the control exists: a routine mints one chat per step, so on a workspace that runs routines the page was already full of them before any client-side filter could look at it, and a person’s newest conversation was not below the fold, it was outside the query. Chat Surface Limits works through the arithmetic. Switching scope therefore re-runs the fan-out, and the count sits on the active tab only — the other scopes’ totals are genuinely not known, and a number we would have to invent is worse than none.
3

Unread, Live, and one agent

Open Filter. Read state, activity and a per-agent list — the same panel, the same facets-AND-together behaviour, the same “Clear all” that every other sidebar on this app has.Unread and Live used to be two thirds of a segmented strip beside All, which was wrong twice over: they are predicates sitting in a control whose third option is a scope, and they read “Unread 0 · Live 0” on most visits, which is how a filter teaches its reader that filters here do nothing. In the panel they compose with the bucket — “unread routines” is a real question, and the exclusive strip could not express it, because choosing Unread threw the scope away — and the count badge on the Filter button is what keeps them from being hidden by the extra click.Click Unread only, then open the row it shows. The count drops immediately and the row stays where it is. Reading something must not delete it from the list you are reading it from, so the open conversation is pinned into the list whatever the toggles say — and the counts are computed without that pin, so “Unread 0” can never sit above a list with a row in it.Live means the agent is working right now. It is driven by the agent.status workspace event rather than by a roster fetched once at mount, so it flips at the moment a run does: send a message and watch it go to 1, then back to 0 when the reply lands. The dot rides the agent’s portrait on the row it belongs to.Collapse (the panel icon at the end of the toolbar) folds the column to a 9px rail with the expand button in it — same control, same place, same behaviour as the other two explorers.Say what did not happen on any of that: no navigation, no refetch, no page transition — picking a conversation is a useState plus a replaceState. Picking used to be a route change, and the whole dashboard chrome tore down and rebuilt every time. Nobody can feel a useState; everybody feels a page transition.
4

A thread that names itself

At the foot of the column, Not started yet collects the agents nobody has talked to. Click it — or the + beside the Conversations heading, which opens the same picker over every agent rather than choosing one for you — and pick casey. That mints a conversation id locally and posts nothing — arriving is not sending, and a conversation you never write in must not leave a row behind.Type a real sentence and send it. Something you will search for in step 5:
Watch the column. The thread appears, and its name is the first message cut to 60 characters at a word boundary, with an ellipsis only if something was actually cut. chats.title had been in the schema since the first migration with nothing writing it, so every thread in every list used to read “Untitled session” — which makes a list of conversations useless the moment there is more than one.Three properties worth saying out loud:
  • It fires once. Rename a thread yourself and nothing overwrites it — not the next message, not a reload.
  • It never blocks the send. The message goes first; the title follows as a separate request. If that request fails, nothing is reported and the session stays untitled, which is where it already was.
  • A message that yields nothing usable leaves the thread untitled. ??? is not a name, and a thread called is worse than one with no name.
Copy the URL now. The session id is the ?session= parameter, and step 6 needs it.

Step 2 — The two kinds of chip

Chips render under an empty conversation, so you need a second, empty one with casey. The column cannot give you that today — Not started yet only lists agents with no conversations at all, and casey now has the one from step 1 — so mint it from the CLI and open its deep link:
Then open /chat/casey?session=<that-id>. A side benefit: the conversation was created by the CLI, so its header carries a CLI origin chip, which is the same chip step 4’s webhook conversation gets as WEBHOOK. After the first turn the chips come back as follow-ups under an assistant reply, three at a time.
1

Show the fallback first

Open /chat/riley in another tab. Four chips: Help me get started · What can you do? · Show me your skills · Run a quick task.That is the built-in role pack, and it is what every unconfigured agent has always shown. Ten seconds on it is worth it, because it is what the next tab replaces.
2

Then the authored ones

Back on casey: three questions somebody wrote for this agent’s actual job, and one chip that is not a question at all.
3

The rule the rail exists to enforce

A chip that opens something must not look like a chip that sends something.
A question chip sends its text the instant it is clicked. A form chip opens a sheet and sends nothing. Somebody who taps Report a bug expecting a form and finds a message already on its way to the agent has been lied to — and it is not recoverable, because the message is gone.So the form chip is marked four ways, and none of them is a colour:Forms lead the rail: somebody sat down and authored them for this agent, and a static question is the cheaper thing. The rail shows six chips under an empty conversation and three as follow-ups; anything past that collapses into +N, which opens the full catalogue upward, out of the rail — the rail sits directly above the composer, and a panel opening downward would cover the input it is offering to fill.

Step 3 — The questionnaire, end to end

1

Open it

Click Report a bug. A sheet grows above the composer, inside the same column, capped at 560px, with the conversation still visible above it.It is deliberately not a centred modal, for two concrete reasons: a centred modal over a chat is a drop target sitting on top of the thing you were dragging from, and on a phone it hides the keyboard-adjacent composer, which is the one piece of UI a phone user needs to stay oriented. Below 900px the same component becomes a 90vh bottom sheet instead.Escape closes it and sends nothing. So does the ✕, so does Cancel, and on mobile so does tapping the scrim.
2

Fill it

  • What went wrongThe credential picker shows no brands on a phone
  • Steps to reproduce — three numbered lines
  • Severity — leave it; the field carries a default and the sheet seeds it
  • First seen — leave it empty on purpose
3

Read the preview

Click Preview message. This is the whole message, verbatim — the rendered template plus the attachment block the composer will append, not just the half the sheet owns.Point at what is missing: there is no First seen: line. An unanswered optional field takes its whole line with it, label included. That is the only rule here you cannot read off the template.Submitting sends an ordinary user message, so the person sending it is entitled to read it first — and an author writing a template meets a broken one here, while writing it, instead of in somebody’s transcript.
4

Answer the photo field, and send

Screenshot or recording is a photo field with its own upload control. Drop any small image on it.Note what the composer does while that control is on screen: it hides its own chip list. There is one attachment list per session, and two views of it read as two attachments — so the sheet wins, because that is where the question is. The chip under the field is the field’s answer, and removing it there is editing that answer.Press Send.
5

Read the provenance line

Above your sent message: a clipboard glyph and via Report a bug.Now reload the page. The badge is still there. That is the part worth demoing rather than describing.Provenance used to be an in-memory map keyed by the rendered message text, and text is not an identity: two identical submissions collided, and a reload lost every entry. What ships now is a submission envelope carried as metadata with the message — the form id, its version, the answers, and which upload answered which field — minted at Send, persisted by the server, and read back off the turn.Two deliberate holes, worth naming before someone finds them: Regenerate re-sends the text without the envelope, because a submission id is minted once per Send and re-sending would record a submission the user never made; and edit-and-resend strips it, because an envelope claiming “this is the text the form rendered” stops being true the moment the text is edited by hand.

Step 4 — An attachment, and what the agent is told about it

Uploading puts bytes in the container. It does not tell the agent they are there. That was the bug, and the fix is the thing to show.
1

Attach from the composer

Drag a file onto the composer, or use the paperclip. Type a caption:
Send it, and read your own turn in the transcript:
The transcript shows that block because it is exactly what the agent received. Four decisions are visible in five lines:
  • It is an ordinary user message. No new envelope, no adapter-specific framing, nothing the agent has to be trained for — which is why it works across every CLI adapter.
  • First person, past tense. Anything imperative — “Read the following file”, <attachments>… — reads as an injected system directive, and an adapter hardened against exactly that shape is entitled to ignore it.
  • The relative path, not the absolute one. The agent’s working directory is /output/<agent-slug>/, so attachments/<chat>/<id>/<file> opens as-is, and the crew slug stays out of the user’s own transcript.
  • One path per line, unquoted. Spaces, quotes and brackets in filenames are common; the line break is the only delimiter none of them can forge.
Then compare with step 3: the image you dropped into the form’s Screenshot field is not in this block. The rendered template already named it under the field that asked, and a file announced twice reads as two files.
2

Send a photo with no caption

Clear the text, attach only a file, press Send. It goes. An attachment is content — a photo needs no caption, and requiring text was how a picture sent from a phone used to disappear without a trace.If that was a brand-new session, look at its name: a message that is only an attachment is named after the file.
3

The camera

Narrow the window below 900px, or open the page on a phone. The composer switches to its mobile branch and a camera button appears beside the paperclip.capture="environment" next to accept="image/*" is what makes a phone open the rear camera instead of the document picker. Without it, the only route to a photo is camera app → gallery → file browser: three screens for the most obvious thing anyone does from a phone. The mobile composer used to be a bare input with no attachments at all, so the one device that actually has a camera was the one surface that could not send a picture.The upload is not a second code path. Camera, paperclip, drop and paste all run the same uploader, with the same size guard, the same failure toast and the same Retry.

Step 5 — ⌘K finds a phrase, not a name

Go back to /chat — which reopens your freshest conversation — and press ⌘K (Ctrl+K elsewhere). Type a phrase from step 1:
This works from inside a conversation too. It did not use to: /chat/<agent> mounts the chat’s own command palette, that palette bound ⌘K as well, and neither listener stopped the other — so one press opened both dialogs stacked. The chat palette is ⌘/ now (or the Commands button in the chat header), and ⌘K means the same thing on every route.
A Conversations group appears: the matched snippet, the agent that said it, and how long ago. Selecting a row opens /chat/<agent-slug>?session=<chat-id> — the same deep link a chat notification uses, which is why the session stays a query parameter and never becomes a path segment. What makes this group different from every other group in the palette:
  • It asks the server on every keystroke — debounced, and each keystroke cancels the request before it. Every other group is fetched once when the palette opens and filtered in the browser. Messages are the largest thing in a workspace and a match is a phrase rather than a name.
  • Two characters is the minimum before anything is sent.
  • A failure renders nothing at all. Nobody explicitly ran this search, so it never raises an error: if the search mirror is not configured on the deployment, the group simply does not appear.
  • The search is from-now-on. Only turns recorded after the mirror was configured are indexed, and there is no backfill. Searching for a phrase you said four minutes ago in step 1 is the demo that always works.
Workspace scope spans at most 400 agents, taking the most recently created ones, and nothing in the response says the scope was cut. If you are demoing anywhere near that number, say so rather than letting it surprise somebody later.

Step 6 — The same surface, without a browser

Three commands. The point of each one is different.
1

Render a form with no browser

This is the inverse of the house rule that every endpoint gets a CLI command: ask forms deliberately have no endpoint of their own. So this is not a wrapper — it is the only way to answer the question an author actually has, which is “what does this template produce?”. It renders through the same package the server uses, pinned to the browser’s renderer by one golden fixture, so what it prints is what the composer would send.No --var first_seen= was given, so that line is gone — the drop-the-whole-line rule from step 3, seen from outside. And note one honest difference from the sheet: the sheet seeds a field’s default, the CLI does not. --var severity=Major is here because the CLI answers only what you tell it.Output is plain text on stdout with no framing, so it pipes into a file, a diff or a review comment.
2

Rename a thread from the terminal

Take the session id out of the URL you copied in step 1, and leave that browser tab open while you run this:
The sidebar row repaints. The new title is broadcast on the workspace channel as chat_renamed, so an open session list in another tab follows without polling.And the row does not move. A rename is not activity, so the server does not bump last_activity_at and the client splices only the title into the row it already holds — fixing a name never shoves a thread to the top of a list ordered by recency.The title in the success message is the stored one, so the server’s normalisation is visible: folded onto one line, control and invisible formatting characters stripped, trimmed.
3

List what is attached

Attach the same-named file twice and the reason for the <attachment-id> segment demonstrates itself: two attachments, two checksums, two paths, neither replacing the other. The filename used to be the last segment and the second upload overwrote the first.The ID column is what crewship chat attachments delete takes, and that delete is the only way to reclaim a chat attachment’s bytes short of deleting the whole chat — chat blobs are not content-addressed, because the path is what the agent is told to open, so they sit outside the background reclaim sweep.--agent skips the lookup. Without it the CLI walks every visible agent asking for its chats until it finds the id.

Step 7 — Four refusals, on purpose

Each is one action, and each shows a different standard.
1

A form that will not submit

Open Report a bug, clear What went wrong, press Send:
Now paste 121 characters into it and press Send again:
Then drop four images on the Screenshot field:
Every refusal names the field by its label. “Something is missing” across five inputs is not a message. The rules live in one module per language, pinned to each other by a shared fixture, so the sheet and the CLI refuse the same answers with the same words:
A preview whose whole promise is “this is what gets sent” must not print a message the console would have refused.The sheet surfaces one problem at a time; the CLI prints them all.
2

A refused upload

Make something over the cap and try to attach it in the browser:
The composer refuses it before anything is uploaded, with a toast reading big.bin exceeds 25 MB — and no chip appears, because nothing was attached. The same bound is enforced again server-side, which is what the CLI meets:
One bound, enforced twice, and both say so.Contrast it with an upload that fails after reaching the network — a provisioned-but-stopped crew answers 409 with the remedy in the message. That one leaves an error chip on screen with a Retry, keeps the bytes in the browser so Retry has something to send, and refuses the Send while it is there: a composer that visibly holds a file and does nothing when you press Send was the original defect. It needs a stopped crew to produce and is not scripted here.
3

A title the server will not store

A refusal, never a silent truncation. Put it beside the auto-derived title from step 1, which is cut — at 60 characters, at a word boundary, with the ellipsis visible in the result. One is a value you chose and the server will not quietly change; the other is a label the product wrote for you, and it shows you what it did.
4

A list the server will not take

The refusal names the offending item by position, and nothing is writtencasey keeps the three questions from step 0. That is the standard the suggested-prompt and ask-form caps all meet: refuse loudly, name what was wrong, write nothing.

What this walkthrough leaves out, and why

Say these rather than letting somebody find them. There is no interactive approval card in chat. Approvals are decided on /approvals. AskUserQuestion renders deliberately inert in the transcript, and a test pins that it must not look clickable. Two event names are reserved for an in-chat decision card so that whoever builds it emits the agreed one — the card itself does not exist. The chip and form funnel is measured, and the measurement goes nowhere. The event vocabulary ships with a bounded in-memory buffer and no network transport at all — no fetch, no beacon, no socket — because PRIVACY.md promises that no usage analytics leave the install. Nothing in the console or the CLI reads that buffer, so there is nothing to put on a screen. See Chat interaction events for the vocabulary and the reasoning; two of its session events are declared and not yet emitted, because their call sites live in files another workstream owns. The one silent cap has no screen affordance. Past twelve agents, the thirteenth agent’s conversations are never fetched, so they are simply absent from the column with no row standing in for them — indistinguishable from an agent nobody has talked to. Opening that agent by name (/chat/<slug>) exempts it from the cap, but browsing the column will not reveal it. Do not demo on such a workspace without saying so first. Chat Surface Limits is the honest list of every bound and whether anything tells you. Two render bounds truncate silently. One substituted answer over 2000 characters, and a finished form message over 32000, are both cut with no warning. A valid form cannot reach the message bound, so in practice this is a very long textarea answer — ask-preview shows the truncated result exactly, which is the only place an author can see it before a user does. Suggested prompts and ask forms are not on the create path. POST /api/v1/agents ignores both keys and returns 201 — the columns are left null. Everything that sets them, including the seed and crewship apply, does it with a follow-up PATCH.

See also