Emmanuel Whyte
AI conversation studio

A-Convo

Year2026
KindAI conversation studio
Built withTypeScript, React, WebSockets, local models
StatusIn progress, kept as an artifact
Runs locallyMulti-agent AI

It started with my own drawings. I wondered what it would be like to hand one to several AI agents and listen to them argue about it: what they would notice, where they would disagree, and whether any of them would see what I meant.

A-Convo is how I found out. You give it something to discuss (an image, a PDF, a clip, or just a question), choose a cast, pick the local model everyone speaks through, and press start. The speakers take turns, argue, agree, change their minds, and a hidden director decides who speaks next and steps in when the talk drifts. A real person can sit in as a live speaker, typing or talking into the thread.

I call it a fun experiment that is really hard to nail down. The interface is busy on purpose: every dial I wanted while chasing a conversation that feels alive is out on the table.

A saved conversation: The Contrarian and The Empath discussing Duchamp's Fountain, with a live status bar for each speaker
A conversation, turns 30 to 32Asked: when did the historical and cultural context take precedence over the voice and intention of the artist? The Contrarian and The Empath end up on Duchamp's Fountain. The rail on the right shows each speaker's state.
The source

A drawing, and a question

Everything begins with what the speakers look at. Drop in an image (here, my drawing Rabbit hole), and a local vision model writes a visual brief that every speaker receives, which I can read, edit, or reroll before anyone talks. Then a framing prompt steers what the conversation is about, and a topic-heat dial sets how heated it gets.

Rabbit hole uploaded as the source, with the visual brief the speakers will receive
Image inThe drawing becomes source memory; its visual brief can be edited before the talk.
The framing prompt asking what the rabbit is guarding, with topic heat, conversation flow, and the local model
Question in“What is the rabbit guarding, and does the loose pastel work against the armour or with it?”
The cast

Choosing who is in the room

Speaker setup: The Contrarian and The Empath with roles, participation, description, preset, tuning, and voice
CharactersEach speaker has a role (AI, live AI, or a real person), a preset, sliders for conviction, wit, empathy, and warmth, and an optional voice.
Conversation Studio setup: input material, workflow templates, and a memory budget for the chosen local model
The studioDrop in source material, start from a template, and see how much graphics memory the chosen local model needs.
  • Host interjections. Any speaker can be switched from participant to host, who guides the room instead of arguing in it.
  • A live status bar. Each speaker's lane lights up when they are thinking, speaking, or next in line.
  • A preferred local model. Pick the model everyone speaks through, with its memory cost shown before you start.
  • Everything kept. Transcripts, prompts, and events are saved in a library you can search, reopen, or fork into a new setup.
Local by design

Voices and a privacy check

Voice Profiles: reusable voice samples kept on this computer only
VoicesReusable voice samples, stored on this computer only.
Local Readiness: health probes, including a privacy boundary showing no external model or telemetry providers active
Local readinessOne probe checks the privacy boundary: no outside model or tracking service is in use.
Honestly

Why it is still in progress

Getting language models to hold a real conversation is much harder than getting one to answer a question. Some of what I ran into:

  • Everyone gets polite. Left alone, speakers drift toward the same agreeable voice. The skeptic starts validating; the empath starts debating.
  • Verbal tics. Openers like “Mm. But…” and “That framing…” repeat, and a “this isn't X, it's Y” habit kept leaking in until I built a guard for it.
  • Attention is not yet human. Real people interrupt, pick up a thread from five minutes ago, and let silence sit. The director's timing still feels like taking turns.
  • Voices and echoes. Spoken audio with a live person in the room brings echo and interruption problems I have not solved, so it runs as text for now.
  • Maybe unnecessary. One good model might answer most of these questions just as well. I keep A-Convo as an artifact of the question, and because watching the characters argue about a drawing is genuinely fun.
Developer notes

What changed lately

  • Speakers were drifting into the same polite voice; each now keeps its stance and avoids repeating its own openers.A-Convo

Tech stack

  • TypeScript, Express, WebSockets - the conversation engine, streamed live
  • React and Jotai - setup, the live view, and the library
  • Open-weight language models - one shared local model for the speakers and the director