What Happens Behind a Three-Button Remote
I recently built a remote control with three buttons.
That sentence makes the project sound more useful than it is.
The actual goal was to lie on the sofa and watch Douyin on a TV.
The TV is connected to an Apple TV. A Mac mini lives in another room and is usually on. My first attempt was the obvious one: run Douyin on an iPhone and AirPlay it to the Apple TV.
It works, but my iPhone 13 gets hot, eventually starts to feel sluggish, and wants to stay on a charger. At one point I seriously considered putting it on a wireless charging stand with extra cooling and turning the phone into a dedicated player.
That immediately created a second problem: if the phone is sitting on a stand, how do I control it from the sofa?
I spent some time going through Apple’s accessibility stack. Switch Control looked promising but was not really designed to turn a Mac into an arbitrary iPhone remote. Voice Control was more interesting. In principle, the Mac could literally play a recorded command such as “swipe up,” the iPhone could hear it, and Voice Control could operate Douyin.
Technically possible. Also ridiculous.
I considered XiaoAi, Xiaomi’s voice assistant, as well. Voice is a natural fit here: say “next,” hit a public endpoint, and let the Mac mini perform the action. The path is probably viable, but maintaining a Worker, authentication, tokens, and a local bridge for three media controls felt excessive, so I left it for later.
Then I noticed something I should have checked much earlier: Douyin has a native macOS app.
That changed the problem completely.
flowchart LR
Mac["Mac mini<br/>Douyin.app"]
ATV["Apple TV"]
TV["Television"]
Mac -->|AirPlay| ATV
ATV --> TV
The Mac mini could now be the player. The iPhone no longer needed to decode video, stay on charge, or remain part of the playback path.
The only problem left was control.
An Apple TV remote does not magically become an input device for the Mac that is mirroring to it. A wireless mouse works, but using a mouse from a sofa to browse short videos feels worse than using a phone.
The actual control surface I needed was this:
Previous
Play / Pause
Next
For a one-off solution, I should have stopped here. A tiny HTTP server on the Mac mini plus a mobile web page would have been enough:
Phone
↓
HTTP / WebSocket
↓
Mac mini
↓
Keyboard Event
↓
Douyin
Instead, I happened to be working on Free4Chat at the same time, so the three buttons ended up sitting on top of a much larger system.
For this remote alone, the tiny HTTP server would still be the better engineering choice. The experiment I actually cared about was different: can an agent integrate a local capability, project only its bounded semantics into a temporary Room, and hand another human a usable interface over it?
I wrote the infrastructure history in Four Evolutions of a WebRTC Chat Room. Free4Chat started as an anonymous WebRTC room. It still has Rooms and realtime communication, but the participants are no longer limited to humans. A local coding agent can join the same Room, take a Task, keep working on its own machine, ask for permission, deliver artifacts, expose a Live View, or publish a generated interactive app.
The Douyin remote became a useful way to test whether those pieces actually fit together.
Put the coding agent in the Room first
I did not start by writing the remote.
I created a Free4Chat Room, attached Codex from the Mac mini as an Agent participant, and created a Task in the Room. The requirement was simple: Douyin is already installed on this Mac; build a way for me to control previous, next, and play/pause from a phone.
The rest of the work happened inside that Task.
A Task is more than one model turn. I can choose the agent/model and execution mode, leave the Task and come back later, steer it while it is working, interrupt it, approve or reject privileged actions, and inspect the outputs it publishes along the way.
The final deliverable does not have to be a chat message. It can be an artifact, a Live View, or an interactive Task App.
A simplified version of the control flow looks like this:
sequenceDiagram
participant H as Human
participant R as Free4Chat Room / Task
participant A as Codex
participant RT as Agent Runtime
participant OS as macOS
H->>R: Create Room
H->>R: Attach Agent
R->>A: Start / resume agent session
H->>R: Create Task
R->>A: Deliver task context
A->>RT: Inspect local environment
RT->>OS: Read / operate local system
OS-->>RT: Result
RT-->>A: Result
A->>R: Permission request
R-->>H: Approve / Reject
H->>R: Approve
R->>A: Continue
A->>RT: Build and register Adapter
A->>R: Publish artifact / Live View / Task App
R-->>H: Interactive result ready
Once an agent runs for minutes or hours instead of producing one answer, model intelligence is only part of the problem.
Something has to own the session. Something has to resume it. Streaming updates need to be correlated with the right Task. Permission requests must return to the right human. Cancellation, interruption, task handoff, artifacts, and generated UI all need lifecycle boundaries.
This is why I find protocols such as ACP, the Agent Client Protocol, interesting. ACP is concerned with the host-to-agent side of the system: creating and resuming sessions, streaming updates, permission requests, cancellation, and agent lifecycle.
MCP sits on a different boundary. It is about how an agent discovers and uses tools.
Very roughly:
flowchart LR
Human["Human / Host"]
Agent["Agent<br/>Codex / Claude"]
Tools["Tools / MCP Servers"]
World["External Systems"]
Human -->|"Session / streaming / permission / cancel<br/>ACP concerns"| Agent
Agent -->|"Tool discovery / invocation<br/>MCP concerns"| Tools
Tools --> World
Free4Chat does not need to claim that every piece of this stack is a new protocol of its own. What matters is that once independent agents become participants in a realtime Room, these problem spaces collide.
There may be multiple humans, multiple agents, several Tasks, tools owned by different runtimes, and UI generated as the result of a Task. Agent-to-agent addressing and handoff starts to resemble the A2A problem space. Turning an agent result into a human-operable interface touches the A2UI problem space.
The remote is small enough to make those boundaries visible without requiring a fake enterprise workflow.
The agent first created a local capability
Codex inspected the macOS Douyin app and found that the required actions could be driven with accessibility and keyboard events. There was no need to reverse engineer Douyin or invent a private API.
But I did not want the runtime contract to be pressArrowDown() or pressSpace().
And I definitely did not want a DouyinIntegration in Free4Chat Core.
The agent created a small local Adapter that exposed semantic actions instead:
douyin_remote
├── previous
├── play_pause
└── next
Conceptually, the descriptor looks something like this; this is intentionally simplified and is not the full current schema:
{
"id": "douyin_remote",
"actions": [
{ "name": "previous" },
{ "name": "play_pause" },
{ "name": "next" }
]
}
The Runtime only needs to understand a generic capability protocol: list, describe, invoke, and, where appropriate, observe.
The Adapter owns the integration details.
flowchart TB
Runtime["Agent Runtime<br/>Generic Capability RPC"]
Adapter["External Adapter<br/>douyin_remote"]
OS["macOS<br/>Accessibility / Keyboard"]
App["Douyin.app"]
Runtime -->|"describe / invoke / observe"| Adapter
Adapter --> OS
OS --> App
This boundary matters more than Douyin itself.
If every local device or desktop app required code in Core, Free4Chat would slowly turn into an integration catalog. Printers, media apps, smart-home devices, serial hardware, vendor SDKs, and credentials would all start leaking inward.
I want the opposite boundary:
The Runtime knows capabilities, not integrations.
The participant owns the integration. The Room only sees a bounded semantic projection of what that participant is willing to share.
I had already tested the same model with an Epson printer. That Adapter talks to local CUPS and exposes a printer status capability. A printer and a media app look unrelated, but from the Runtime’s point of view they are just different semantic capabilities:
printer_status
douyin_remote
One reads local device state. The other controls a local application.
The Room does not need to know how either one works.
Then the agent built the remote itself
At this point Codex could already control Douyin from the Task.
I could type “next,” the agent could invoke douyin_remote.next, and the video on the TV would change.
That is still a terrible remote control.
So the Task continued. Codex generated a small interactive Task App against the capability it had just registered.
Full screen on the phone, it was basically this:
┌────────────────────┐
│ │
│ Previous │
│ │
│ Play / Pause │
│ │
│ Next │
│ │
└────────────────────┘
That is the entire user interface.
I joined the same Room from an iPhone, opened the Task App, pressed Next, and the video on the TV changed. In use it feels like an ordinary remote; there is no pause while Codex thinks. I later joined the same Room from a second Human device as well. The same Task App could control the Mac mini, and its shared state stayed synchronized between the two clients.
At that point the original requirement was done: three buttons on a phone, one action on the TV.
If the story ended here, this would just be another “an AI wrote a web page” demo. Follow the Next button downward, though, and the page turns into something much larger.
What actually happens after pressing Next
The complete path is closer to this:
flowchart TB
subgraph Phone["iPhone / Browser"]
App["Generated Task App<br/>Previous / Play / Next"]
Host["Room App Host"]
App --> Host
end
subgraph Room["Free4Chat Room"]
Presence["Presence / Identity"]
Task["Task State"]
Permission["Authorization"]
Projection["Capability Projection"]
AppState["App Metadata / Shared State"]
end
subgraph Mac["Mac mini / Agent Participant"]
Runtime["Agent Runtime"]
Router["Capability Router"]
Adapter["douyin_remote Adapter"]
Accessibility["macOS Accessibility / Keyboard"]
Douyin["Douyin.app"]
Runtime --> Router
Router --> Adapter
Adapter --> Accessibility
Accessibility --> Douyin
end
Host -. "Room state / authorization" .-> Room
Host == "Reliable WebRTC DataChannel" ==> Runtime
Douyin -->|AirPlay| ATV["Apple TV"]
ATV --> TV["Television"]
The first thing to notice is that the Room is not a generic remote-control proxy.
There are two different paths in the system.
Control plane and participant data plane
The Room has to coordinate realtime collaboration. It needs to know who is present, which Task belongs to which participant, what an agent is doing, which generated app belongs to the Task, what bounded capabilities an agent currently projects, and what the current authorization state is.
That is control-plane state.
Browsers maintain Room state through the application signaling path, with a Durable Object coordinating the Room lifecycle.
A local capability invocation is different.
If every Next call had to travel through:
Browser
→ Worker
→ Durable Object
→ Agent
→ Runtime
→ Adapter
then the Room would gradually become a universal RPC proxy. Local data would always detour through the cloud. Latency and Worker cost would rise. More importantly, the application core would begin owning integration traffic that belongs to the participant.
So Free4Chat separates the two paths:
flowchart TB
subgraph CP["Control Plane"]
WS["Room WebSocket"]
DO["Room Durable Object"]
Presence["Presence / Identity"]
Tasks["Task Lifecycle"]
Perm["Permissions"]
Apps["App Metadata / Shared State"]
Cap["Capability Projection"]
WS --> DO
DO --> Presence
DO --> Tasks
DO --> Perm
DO --> Apps
DO --> Cap
end
subgraph DP["Participant Data Plane"]
Browser["Human Browser"]
DC["Reliable WebRTC DataChannel"]
Agent["Agent Participant"]
Runtime["Agent Runtime"]
Adapter["Local Adapter"]
Browser --> DC
DC --> Agent
Agent --> Runtime
Runtime --> Adapter
end
CP -. "identity / addressing / authorization" .-> DP
“Participant-direct” here does not mean that there is always a naive browser-to-Mac P2P connection with no infrastructure involved. The path still uses Free4Chat’s WebRTC participant transport and Cloudflare Realtime SFU where appropriate.
It means that capability traffic belongs to the Human participant and the target Agent participant. It does not need the Room Durable Object to relay every application-level operation.
That distinction barely matters for three buttons. It matters much more once the capability becomes a live sensor, a hardware control surface, continuous state synchronization, or many participants talking at once.
Free4Chat originally used WebRTC because humans needed realtime media. Once agents became participants, reliable DataChannels became useful for a second kind of realtime traffic: machine data between participant runtimes.
The LLM is not in the runtime hot path
Another design choice is easier to see with a remote control than with a normal agent demo.
Pressing Next should not call Codex again.
This would be a bad runtime:
Human Click
↓
LLM
↓
Interpret "Next"
↓
Choose tool
↓
Capability invoke
There is no value in paying model latency and accepting model uncertainty for a three-state interface.
The task has two phases instead:
flowchart TB
subgraph Build["Build Time - Agent involved"]
Req["Human Requirement"]
Agent["Codex"]
Inspect["Inspect Environment"]
BuildAdapter["Build Adapter"]
Register["Register Capability"]
Generate["Generate Task App"]
Req --> Agent
Agent --> Inspect
Agent --> BuildAdapter
BuildAdapter --> Register
Agent --> Generate
end
subgraph Run["Run Time - deterministic hot path"]
Human["Human Click"]
App["Task App"]
Host["Room App Host"]
DC["Reliable DataChannel"]
Runtime["Agent Runtime"]
Adapter["douyin_remote"]
Action["Accessibility / Keyboard"]
Human --> App
App --> Host
Host --> DC
DC --> Runtime
Runtime --> Adapter
Adapter --> Action
end
Register --> Runtime
Generate --> App
Codex is the builder.
Once the Adapter and UI exist, the LLM leaves the hot path. Every button press is a normal deterministic RPC.
This is a pattern I increasingly prefer in agent systems: let the model handle fuzzy, high-level, low-frequency work; let normal software handle frequent, bounded, testable execution.
The final artifact of the model is not another prompt. It is software.
Why is it safe to run an app generated by an agent?
Generating the HTML is the easy part.
The harder question is why I should trust an agent-generated page at all.
If generated JavaScript could reach internal Room objects, local runtime credentials, arbitrary network endpoints, or system capabilities directly, a more powerful Generated App would simply create a larger security problem.
A Task App therefore runs through a bounded App Host contract rather than receiving arbitrary ambient authority.
flowchart TB
Generated["Agent Generated App"]
Sandbox["Sandboxed App Surface"]
Host["Room App Host"]
State["Bounded Shared State API"]
Cap["Bounded Capability API"]
Context["Room Context"]
Generated --> Sandbox
Sandbox --> Host
Host --> State
Host --> Cap
Host --> Context
The app does not need to know how the Adapter process starts. It does not receive macOS accessibility privileges. It sees a narrow host API for things such as shared Room state and invocation of capabilities already projected and authorized by the system.
For actions that create a real-world side effect, invocation also stays tied to trusted human interaction. The agent may have generated the button, but that does not mean its page can silently loop over side effects in the background.
For next, this may look excessive.
Replace next with open_door, move_robot, or delete_file, and the same boundary stops looking theoretical.
For generated UI, writing the <button> is the easy part. The harder part is letting that interface participate in a real system without inheriting the authority of everything behind it.
Text, artifacts, Live View, and Task Apps
The remote also made another product boundary clearer to me: what should an agent actually return at the end of a Task?
Chat products naturally make every result look like a message, but a long-running Task has several useful output shapes.
flowchart LR
Agent["Agent Task"]
Text["Text"]
Artifact["Artifact"]
Live["Live View"]
App["Generated Task App"]
Human["Human"]
Agent --> Text --> Human
Agent --> Artifact --> Human
Agent --> Live --> Human
Agent --> App --> Human
Text says: here is the answer.
An artifact says: here is the thing I produced.
A Live View says: here is something worth watching while the task is still running.
A Task App says: here is an interactive result you can continue to operate.
The remote clearly belongs to the last category.
This is the part of the A2UI direction I find more interesting than simple UI generation. If the page is disconnected from the Task, permissions, realtime state, and capabilities that produced it, then it is just generated frontend code.
Once the UI is a bounded human surface over a live agent task and a real participant capability, it becomes part of the agent system itself.
ACP, MCP, A2A, and A2UI meet in the same Room
ACP and MCP have already appeared earlier in the path. Zoom out one level and a Room with multiple humans and agents also starts to touch the problems usually discussed under A2A and A2UI.
What matters more to me than the four labels are the boundaries they describe: host-to-agent control, agent-to-tool access, agent-to-agent collaboration, and agent-to-human interface. In a real realtime Room those boundaries stop looking like separate product categories.
flowchart TB
Room["Free4Chat Room<br/>temporary collaboration boundary"]
HumanA["Human A"]
HumanB["Human B"]
AgentA["Codex"]
AgentB["Another Agent"]
MCP["MCP / Tools"]
Runtime["Participant Runtime"]
Adapter["Local Adapter"]
UI["Live View / Generated App"]
HumanA --> Room
HumanB --> Room
Room -->|"Agent session / permission / interrupt<br/>ACP problem space"| AgentA
Room -->|"Agent lifecycle"| AgentB
AgentA <-->|"Agent collaboration / handoff<br/>A2A problem space"| AgentB
AgentA -->|"Tool access"| MCP
AgentA --> Runtime
Runtime --> Adapter
AgentA -->|"Interactive result<br/>A2UI problem space"| UI
UI --> HumanA
UI --> HumanB
I use “problem space” deliberately. Free4Chat does not need to claim that every arrow in this diagram is a complete implementation of every external standard.
The useful observation is that these are no longer separate product categories once humans and independent agents share the same realtime environment.
An agent has to be controlled. It has tools. It may collaborate with another agent. It may produce a UI. A human may leave and return later. Permissions still need a human boundary. Results need Task identity. Realtime transport has to survive independently of model turns.
At that point, calling the system an “AI chat room” is no longer very accurate.
The Room is a temporary trust boundary
The finished remote also had a useful side effect: it was not tied to my iPhone.
Another human in the same Room can open the same Task App and, if authorized, use the same capability.
That does not copy douyin_remote onto their phone. It does not deploy the Adapter to their device. It does not turn the Mac mini into a permanent public API.
Ownership stays split:
Capability → Mac mini participant
Adapter → Mac mini / Agent
Task App → Task
Authorization → Room
Human → temporary user of the capability
The local capability remains local. The participant projects a bounded description into the Room. The Room scopes who may use it during the collaboration.
flowchart LR
Mac["Mac mini Participant<br/>owns capability"]
Agent["Agent<br/>projects capability"]
Room["Room<br/>scopes trust"]
App["Task App<br/>provides UI"]
Human["Human<br/>temporarily uses"]
Mac --> Agent
Agent --> Room
Room --> App
App --> Human
When the Task ends, the participant leaves, or the Room expires, the temporary collaboration relationship disappears. The local integration and its credentials never needed to migrate into Free4Chat Core.
For a media remote this is mostly convenient.
For an internal service, a printer, home automation, a camera, local files, or hardware that should never become a public cloud integration, the same boundary matters much more.
Three buttons, many state machines
The UI state for the remote is almost embarrassingly small.
It is either usable or unavailable.
The state underneath it is not.
A successful button press may depend on all of these being in the right generation and lifecycle at the same time:
Room membership
Human session
Agent participant session
Agent harness session
Task lifecycle
Task execution state
Permission state
Generated App lifecycle
Room App Host state
Shared App state
Room WebSocket
WebRTC PeerConnection
Publisher / Subscriber session
Reliable DataChannel
Capability projection
Capability authorization
Runtime transport
Adapter process
macOS Accessibility permission
Douyin application state
AirPlay session
Those states do not even live in the same process.
The browser owns some. A Room Durable Object owns some. The SFU and participant transport own some. The local runtime has its own lifecycle. The Adapter is local. macOS owns accessibility permission. AirPlay owns the final display path.
A simplified dependency chain looks like this:
flowchart TB
Join["Human joined Room"]
Task["Task / App available"]
Agent["Agent participant ready"]
Projection["Capability projected"]
Host["App Host mounted"]
RTC["WebRTC participant transport ready"]
DC["Reliable DataChannel ready"]
Auth["Capability authorized"]
Runtime["Runtime route ready"]
Adapter["Adapter reachable"]
OS["OS permission valid"]
Action["Local action"]
TV["TV picture changes"]
Join --> Task
Task --> Agent
Agent --> Projection
Projection --> Host
Host --> RTC
RTC --> DC
DC --> Auth
Auth --> Runtime
Runtime --> Adapter
Adapter --> OS
OS --> Action
Action --> TV
If any precondition disappears, the human may see nothing more informative than “unavailable.”
That is where most of the difficulty comes from. It is not the number of source files; splitting a large hook into five smaller files does not remove any of these lifecycles.
The important part is ownership: which component owns a state transition, which generation is current, who is responsible for recovery, and what evidence tells you that the transition happened.
Free4Chat already had the usual WebRTC state: ICE, PeerConnection, publishing, subscribing, signaling. Adding resident agents brought more independent lifecycles with it: harness sessions, Tasks, permissions, capability projection, and generated applications.
That is the machinery hidden under the three buttons.
Accessibility comes back into the story
The early experiments with Switch Control and Voice Control turned out to be more relevant than I expected.
Apple has spent years making software operable by people who cannot use the conventional mouse-keyboard-touch model. Accessibility trees, alternate input, voice control, keyboard navigation, and system automation all exist primarily for human accessibility.
Agents have become an unexpected second user of that infrastructure.
An agent also has no hand.
A browser has the DOM. A terminal has a shell. A well-designed service has an API, perhaps exposed through MCP.
A desktop application often has none of those.
If a system still needs to understand what is on screen and trigger actions in a GUI, the accessibility layer is one of the few existing machine-readable interfaces available.
It also helps explain how computer-use agents can operate Finder, browsers, IDEs, and older desktop applications that were never designed for automation.
The Douyin Adapter is a modest example. It does not need advanced visual reasoning; keyboard events are enough. But the path started with accessibility features built for humans and ended with an agent using the same substrate to turn a GUI-only application into a semantic capability.
That is a fairly strange reuse of decades of accessibility engineering.
Douyin is not the important part
Once this worked, Bilibili was the obvious next example.
The TV client does not provide the same experience as the normal desktop/mobile product, while the Mac can already run the desktop app or website. If the Mac mini remains the player and AirPlay remains the display path, the phone can still be nothing more than a remote.
Only the capability changes:
douyin_remote
├── previous
├── play_pause
└── next
bilibili_remote
├── previous
├── play_pause
├── next
├── fullscreen
└── favorite
Put that next to the earlier printer experiment:
printer_status
├── state
└── accepting_jobs
What gets reused is not the implementation of “next video” but the whole path:
Local Integration
↓
Semantic Capability
↓
Participant Runtime
↓
Temporary Room
↓
Agent-generated Human Interface
The same shape could apply to a local printer, an ESP32, Home Assistant, a desktop application, an internal company service, or something that only exists on a private network.
The agent can live next to the capability instead of forcing the capability to become a permanent cloud integration first.
The Room only exists when collaboration is needed.
Back to Next
After all of that, the screen still has three buttons.
┌────────────────────┐
│ │
│ Previous │
│ │
│ Play / Pause │
│ │
│ Next │
│ │
└────────────────────┘
I press Next.
The Generated Task App calls a bounded host API. The Room has already established the human, Task, target Agent, and authorization relationship. The browser sends a capability invocation over the reliable participant DataChannel. The Runtime routes it to douyin_remote. The Adapter turns the semantic action into local macOS input. Douyin switches videos. The Mac continues mirroring through AirPlay to the Apple TV.
The picture on the TV changes.
There is no LLM call in that path.
If I had only wanted those three buttons, an HTTP server would have been the better engineering decision.
But Free4Chat already had Rooms, realtime participant transport, long-running Agent Tasks, permissions, generated apps, and participant-owned local capabilities. This silly remote happened to exercise nearly all of them in one small scenario.
When I use it from the sofa, none of that is visible.
It is still just three buttons.