A-Convo
It started with my own drawings. I wondered what it would be like to hand one to several AI agents and listen to them argue about it: what they would notice, where they would disagree, and whether any of them would see what I meant.
A-Convo is how I found out. You give it something to discuss (an image, a PDF, a clip, or just a question), choose a cast, pick the local model everyone speaks through, and press start. The speakers take turns, argue, agree, change their minds, and a hidden director decides who speaks next and steps in when the talk drifts. A real person can sit in as a live speaker, typing or talking into the thread.
I call it a fun experiment that is really hard to nail down. The interface is busy on purpose: every dial I wanted while chasing a conversation that feels alive is out on the table.

A drawing, and a question
Everything begins with what the speakers look at. Drop in an image (here, my drawing Rabbit hole), and a local vision model writes a visual brief that every speaker receives, which I can read, edit, or reroll before anyone talks. Then a framing prompt steers what the conversation is about, and a topic-heat dial sets how heated it gets.


Choosing who is in the room


- Host interjections. Any speaker can be switched from participant to host, who guides the room instead of arguing in it.
- A live status bar. Each speaker's lane lights up when they are thinking, speaking, or next in line.
- A preferred local model. Pick the model everyone speaks through, with its memory cost shown before you start.
- Everything kept. Transcripts, prompts, and events are saved in a library you can search, reopen, or fork into a new setup.
Voices and a privacy check


Why it is still in progress
Getting language models to hold a real conversation is much harder than getting one to answer a question. Some of what I ran into:
- Everyone gets polite. Left alone, speakers drift toward the same agreeable voice. The skeptic starts validating; the empath starts debating.
- Verbal tics. Openers like “Mm. But…” and “That framing…” repeat, and a “this isn't X, it's Y” habit kept leaking in until I built a guard for it.
- Attention is not yet human. Real people interrupt, pick up a thread from five minutes ago, and let silence sit. The director's timing still feels like taking turns.
- Voices and echoes. Spoken audio with a live person in the room brings echo and interruption problems I have not solved, so it runs as text for now.
- Maybe unnecessary. One good model might answer most of these questions just as well. I keep A-Convo as an artifact of the question, and because watching the characters argue about a drawing is genuinely fun.
What changed lately
- Speakers were drifting into the same polite voice; each now keeps its stance and avoids repeating its own openers.A-Convo
Tech stack
- TypeScript, Express, WebSockets - the conversation engine, streamed live
- React and Jotai - setup, the live view, and the library
- Open-weight language models - one shared local model for the speakers and the director
