Vance Sereda GitHub LinkedIn vancesereda@gmail.com

Here's some of what I'm working on

This is an app I made for the coding agent I use, OMP, which is a fork of the Pi coding agent. The phone is driving a session on my computer, and the model is Qwen running on GPUs in my house. No cloud model is involved.

My app → my Pi extension → OMP on my computer → Qwen on my GPUs

The recording

One prompt from the app to Qwen3.8-Flash-Next on my rig, about 94 tok/s. Recorded on an Android emulator against a separate test server, only trimmed and captioned.

An AI agent drove this demo. I told it what I wanted, and it set up the test server and the emulator, logged the app in, typed the prompt, recorded the screen and built this page.

How it works

I took the app shell from Happy Coder, an open source Claude Code client, and threw out what it talked to. Its backend is now a Pi extension I wrote.

  • On my computer a small wrapper runs OMP in a terminal and keeps a connection to a relay server I host.
  • The extension runs inside OMP and talks to the wrapper over a local socket.
  • Prompts from the phone go into the session with pi.sendUserMessage.
  • OMP's own events (text, reasoning, tool calls, subagents) go back to the phone, so nothing is read off the terminal.
  • Sessions I start from the phone run in herdr on my computer, so I can attach to the same session at my desk and keep going.
  • Sessions I start on the computer show up on the phone right away. It's the same session either way, so I can switch between them whenever.

The rig

GPUs2x RTX 3090, 2x CMP 170HX
CPU / boardEPYC 7452, ASRockRack ROMED8-2T
ServingvLLM, model split across all four cards, MTP speculative decoding, 256K context
Qwen FP890 to 100 tok/s
Qwen 4 bitabout 233 tok/s on two cards
500K token input145 tok/s after 95 s to first token