The recording
One prompt from the app to Qwen3.8-Flash-Next on my rig, about 94 tok/s. Recorded on an Android emulator against a separate test server, only trimmed and captioned.
An AI agent drove this demo. I told it what I wanted, and it set up the test server and the emulator, logged the app in, typed the prompt, recorded the screen and built this page.
How it works
I took the app shell from Happy Coder, an open source Claude Code client, and threw out what it talked to. Its backend is now a Pi extension I wrote.
- On my computer a small wrapper runs OMP in a terminal and keeps a connection to a relay server I host.
- The extension runs inside OMP and talks to the wrapper over a local socket.
- Prompts from the phone go into the session with
pi.sendUserMessage. - OMP's own events (text, reasoning, tool calls, subagents) go back to the phone, so nothing is read off the terminal.
- Sessions I start from the phone run in herdr on my computer, so I can attach to the same session at my desk and keep going.
- Sessions I start on the computer show up on the phone right away. It's the same session either way, so I can switch between them whenever.
The rig
| GPUs | 2x RTX 3090, 2x CMP 170HX |
|---|---|
| CPU / board | EPYC 7452, ASRockRack ROMED8-2T |
| Serving | vLLM, model split across all four cards, MTP speculative decoding, 256K context |
| Qwen FP8 | 90 to 100 tok/s |
| Qwen 4 bit | about 233 tok/s on two cards |
| 500K token input | 145 tok/s after 95 s to first token |