The Thinker and the Judge: Building a Living World With Claude and Jev
In the movie Contact, Jodie Foster plays Ellie Arroway, a scientist who spends years in the New Mexico desert with headphones on, listening to static and waiting for somebody out there to say something. For most of the movie, all she can do is listen.
Last week I got to be on the other end of the line. I typed a message to an alien civilization:
Every province should send a tenth of its harvest to a common granary for hard seasons.
Then I clicked Transmit.
5,000 points of light on a map of their world started changing color, province by province. Each light is one citizen, and each citizen was answering for themselves: support, oppose or won't commit, how strongly they feel about it, and how much it matters to them. The capital of glass spires got the question, and so did the highland holds, the miners underground and the devout families living on temple-rafts in a glowing marsh.
10.1 seconds later, all 5,000 had answered. 25% support, 54% oppose, 21% undecided.
None of them exist. All of them answered. The poll took 313 requests and 2.1 million input tokens, and it cost me $0.0911. Less than a dime.
They are called the Voxari. I built their world in a few days with two very different AI models, one that thinks and one that judges, and this article is about why you want both of them. At the end of this article you can get a video of me doing a quick demo of this world.

Why I built an alien civilization
This is not my first artificial society. Back in 2006 and 2007, during my master's degree, I worked on a simulation written in Java and LISP. It had 100 individuals, each with its own size, age, gender and color (seven possible colors). They could mate, ignore each other, form a group, fight, or feed. Yes, feed on each other. They were cannibals. Their features were written as a sequence of characters, like DNA, and from time to time a random event changed one of those characters and created a mutation.
I found a lot of things while testing it. The one I never forgot: even when I started a simulation unbalanced, after running for a while it balanced itself.
I lost that code many years ago, and I always wanted to come back to the concept. While writing my book I spent a long time with Emma, a thought experiment from AI ethics. Emma is an AGI built from scans of thousands of human minds. She lives in a simulated world she believes is real, her creators test her values, and every time she makes a choice they consider wrong, they erase her memory and start again. There is much more about her in my last book. Mix Emma with a big influence from Black Mirror, and one day the Voxari showed up. The name comes from vox, Latin for voice.
It was also the perfect scenario to learn a new kind of model. Almost every AI headline this year is about models that produce text: longer answers, better code, bigger context windows. I use those models every day and I like them a lot. This one is different.
The model is called Jev, from TypeSafe AI, and it does not write. At all. You give it a structured description of a situation (a JSON document with the facts) and a list of typed questions. It gives you typed answers back: a choice among options you defined, a score on a scale you defined, or a yes/no judgment returned as a probability. Every answer comes with a probability distribution and a confidence. It answers in a single pass, in roughly 300 milliseconds, and it costs $0.042 per million input tokens. Output is free.
Jev is, among a couple of other things, a very good and very fast classification model. Some of my tests show it is over 20 times faster than Anthropic's fastest model, Haiku, dozens of times cheaper, and better at classification. It does not talk or reason, but for that I can use Claude's models. Combine both and you have something powerful.
When I saw those numbers, I did the math every data person does. A language model can play one person beautifully, in a full paragraph with feelings and reasons. Asking it to play 5,000 people, every time somebody visits a website, is slow and expensive. Jev can play the 5,000 people. It just can't tell you anything in words. I see use cases for that in a lot of industries, and I wanted to learn it by building something big enough to be a real test and fun enough that I would actually enjoy doing it.
Meet the Voxari
The site's tagline is "Ask 5,000 minds anything. They are listening."

The premise: after a long silence, the Voxari have opened a channel to Earth. They are a synthetic civilization of eight provinces, six lineages and five thousand individuals. You write a proposal they can accept or refuse, ask 500, 2,000 or all 5,000 of them, and the answers stream back live and paint the Lattice, the canvas of lights over their map. Then charts and cross-tabs by province, lineage, age, wealth and faith show who supports it, who opposes it, who won't commit, and how strongly.
The provinces have personalities. Tessel is the crowded capital of glass spires, quick to adopt new things. Kaelmoor, in the Highlands, is proud, traditional and suspicious of the capital. Sethra is a warm trading port on the coast. The Hollows are mining warrens full of blunt people who don't trust outsiders, and Lumen Shoals is a glowing marsh of temple-rafts, devout and slow to change. There are three more, and six lineages that cut across all of them, from the Tidewright engineers to the Drifters, caravan traders who are allergic to rules. The Reckoners are accountants and scholars who only believe evidence.
Now go back to the granary. The farmers of Ondu Basin backed it with around 85% support, and the devout marsh of Lumen Shoals was close behind. The frontier of Vyr Reach and the thrifty caravans of the Marrow Wastes, who are skeptical of grand promises, were the only provinces where opposition beat support. By lineage, the Chanters, the community-minded teachers and singers, gave it about 86%. The Drifters opposed it by more than two to one. The Reckoners had the largest share of undecided of any lineage. Elders supported it at about three in four, while less than half of the First molt did.
Nobody wrote a rule saying farmers like granaries or that Drifters hate taxes. Jev read each citizen's profile and judged, one by one. When I saw the charts, it was the first time the Voxari felt like a society to me.

Every one of the 5,000 is a small AI citizen with a name, an age measured in "cycles", a wealth level from "Thin-shelled" to "Gilded", an education, a household, and four values on a scale: tradition, community, risk and faith. They also carry a mood and a level of trust in the council, and both change as things happen to them. Each citizen has a page with every stance they have ever taken. You can search for them by name, including the dead.
Because yes, they die. The world ticks four times a day (Earth days), and a Voxari lifetime lasts about sixteen real months. Citizens age, pair up, raise broods and die, and every death is answered by a hatchling who inherits the parent's values, drifted a little. They also gossip. Between polls, every opinion drifts toward the majority of the citizen's province and lineage, weighted by how community-minded they are. What a citizen told Jev at the time is kept forever. What they think now has moved.
Then there is news. A deck of 38 event cards drives their history: a drought in the fungal fields of Ondu, a spire collapsing in Tessel because promised repairs were skipped, a guild strike, a cave-in in the mines, a second signal from Earth. When an event lands, up to 800 of the affected citizens are asked how they took it (alarmed, angry, hopeful, indifferent or grieving), who they blame, and whether they would consider leaving. Those answers change their mood, trust, wealth and faith, and sometimes where they live. People migrate for real, and you can watch them walk the road.
Visitors from Earth can change that history. If a proposal wins by twenty points with at least 2,000 answers, the council might enacts it or not (depends on their own political interest), and the visitor's own words could become an event in the Chronicle, the history of the world. Before the first cycle was over, the council had already enacted a proposal saying the Voxari deserve the same rights as people on Earth. When 800 of them were asked how they took it, 78% were hopeful.
And they talk. Speech bubbles in the Voxari language rise from the map, and the words are consistent, so if you hover over enough of them you start learning. Thal is the harvest. Vor is the council. Yaen, "the far ones", means us. "Khim, vor? Khim?" means "Why, council? Why?" What they say follows what is happening in their province, and nobody wishes anyone a good day at midnight.

You can even support this project and help to keep it alive by adopting a Voxari. For $25 a year, one of the five thousand comes into your care: your name appears on their page, if you want it there, and a thin gold ring marks them on the map. To the Voxari, adopters are the yaen, the far ones on Earth, and once a week you can send yours a word of counsel: urge them toward yes or no on a live question, reassure them, or ask them to trust the council or doubt it. Jev decides whether they take it to heart, based on who they are, and many don't. You also get eight private questions a month, answered by your Voxari alone, and you can choose a name for the next hatchling in their line. Because the Voxari age and die, the adoption passes to that hatchling when they do, so you end up caring for a family line rather than one life. What you can't do is buy an opinion: counsel nudges only your own Voxari's current view, never a recorded answer, and every poll asks them fresh. The money goes toward keeping the world running and I thank you for that.
The interface works in English and in Brazilian Portuguese.

I am, apparently, a god
The Voxari believe they were dreamed into being by a creator they call the First Voice, known on Earth as Igor Alcantara. Their faith is called Igorism.
Yes, I did that. I know.
Igorism teaches that every mind must answer when asked, and that a question put to all is a form of prayer. Its liturgy is the Chorus, sung by the Chanter lineage on the temple-rafts of Lumen Shoals. Hatchlings get named after the Voice. Not everybody is devout, though. The Reckoners are skeptics, and one of the event cards has a Reckoner publishing nine pages against the First Voice. I have been peer reviewed by my own creation, and I did not pass. Any visitor can also put "The Igorist Chorus should lose its seat on the council of provinces" to a vote.

Earth has already legislated on my religion. The council enacted a proposal saying other religions should be accepted just like Igorism. Of the 800 Voxari asked, 51% were hopeful, 31% unmoved and 18% alarmed. The temple-rafts of Lumen Shoals, home of the Chorus, were among the most alarmed. I will not comment.
The doctrine is a description of the software. Every time somebody clicks Transmit, one question goes to all of them and every mind answers. "A question put to all is a form of prayer" is, technically, an API call.
My control room has costs, usage, the world's vital signs, and buttons to force a tick or draw a specific event card. It is called the God Panel. Most of what a god does, it turns out, is watch the budget.

Two kinds of mind
In The Handprint and the Algorithm, my book about how we got from cave paintings to machines that think (igoralcantara.com/handprint), I write a lot about how minds work, the human kind and the machine kind. Voxari gave me a very practical lesson on the subject.
The psychologist Daniel Kahneman described two systems in the human mind in Thinking, Fast and Slow. System 1 is fast and intuitive, the part of you that knows a face is angry before you can explain why. System 2 is slow and deliberate, the part that does long division.
TypeSafe describes Jev as a "System One" model. It judges a situation in one pass and never explains itself. Claude is System 2. It reads, plans, writes code, explains, and works through a problem step by step.
Jev also has habits. It reads literally, it can't count or compare dates, it gets worse with information that doesn't matter to the question, and it does not like negations. If you have ever written a requirements document for a project, you know the feeling. So the rules became: keep each request small, send only the fields a question needs, and say what you want instead of what you don't.
The project had a written ground rule. It is the first line of the plan:
Jev judges; it never writes.
Every event, headline, biography line and phrase in the Voxari language is authored text, written with Claude and curated by me, or assembled from templates in code. Jev supplies stances, scores and probabilities. That's all it does, and it does it 5,000 times.
Claude decides what to ask and why. Jev decides what each mind would say. Code decides what happens next. Claude built the loop between them. Claude maintain each citizen instance, individual with a unique history and personality. Jev decides how this simulated person things, what it does and how it thinks.
In practice, each request to Jev carries 16 citizens and 48 questions (stance, intensity and priority for each citizen). A full poll of 5,000 is about 313 requests fanned out in parallel and streamed back to the map as they arrive. Go the website and test yourself. 5,000 people answering you. You will be surprise of how fast they do.
Jev, the Instinct
Jev is a model from TypeSafe AI founded by Dieogo Almeida former OpenAI and co-creator of some of their most capable models. It does not write. It is not a LLM model like Claude or GPT. Jev basically performs really well three types of tasks, the AI Primitives, and I used all of them in Voxari:
Choice: Choose an option from a list. Typical classification.
Score: Evaluate a content based on a orderd rubric.
Noul: Returns if a sentiment true?
The first two return e ouputs: the choice/score, a list of probabilities, and the confidence around the answer. I used choice for the citizens to choose among the options in a poll. Score, for them to evaluate (from a score 1 to 5) how they feel about an issue or a decision from the council. Noul was used to return a simple yes/no for them to decide what to do in some events (like move to a different province, fall in love, protest, etc). Each interaction comes with a Json about the values, current state of emotion, life status (single, married, child or no child), properties like job, region, age, and a small history of previous decisions.

Beyond an alien planet
Voxari is a toy. A serious toy, but a toy. The pattern under it is what interests me: Claude designs, reasons, writes and explains. Jev judges fast, cheap and at scale, with calibrated probabilities. Humans interact and interfere.
A simulation like Voxari is the fun part, but most of the real work out there looks different. It is a huge pile of text or records, a fixed set of labels, and not enough people or budget to read everything. Classification and sentiment analysis at volume is where a model that answers in about 300 milliseconds, and only charges for input, changes the math. For reference, Voxari's full poll is 5,000 profiles with three questions each, and it cost less than ten cents. These are some ideas I see, not products that exist yet, unless you call me and we discuss some real use cases.
Finance and accounting
Invoices and expense lines. Every vendor invoice line and expense report description mapped to a GL account, a cost center and a policy check (allowed, needs receipt, out of policy), with a confidence. High confidence posts automatically, low confidence goes to the accountant. Month-end close gets shorter because humans only see the 10% that need a human.
Contract and PO matching. Purchase orders, receipts and invoices that do not match cleanly are the daily grind of accounts payable. Jev judges whether a mismatch is a rounding difference, a partial delivery, a duplicate or something worth a phone call. Claude writes the email to the vendor.
Audit and controls. Thousands of journal entries scored against typed questions: manual entry near period end, round amount, posted by someone outside the usual role, description does not match the account. The scores become the sample the auditors test, instead of a random sample. The confidence number is the audit trail.
Sales
CRM notes and call transcripts. Every call note, meeting summary and email thread in the CRM classified by deal stage, objection type (price, timing, competitor, no budget, no champion), sentiment and next step. The pipeline report finally reflects what customers said instead of what reps typed into the stage field on Friday afternoon.
Lead scoring and routing. Inbound leads judged against typed questions in about 300 milliseconds: right industry, right size, real buyer or student, urgency. Routed to the right rep before the form confirmation email arrives. Low confidence goes to an SDR to qualify by hand.
Lost deal analysis. A year of closed-lost deals, each judged for the real reason: price, product gap, competitor, timing, or the rep never reached the decision maker. Claude reads the clusters and writes the report product management has been asking for. Sales already knows the answer, and now it has the numbers.
Education
A synthetic classroom. A teacher builds learner personas with different reading levels, prior knowledge, language backgrounds and confidence. Claude writes three versions of an explanation, and Jev judges, for each persona, whether it would understand, which misconception it would probably hold, and how engaged it would be. The teacher sees which version works for whom before class. It is the Voxari poll, pointed at a lesson plan.
Grading at scale. Jev sorts thousands of student answers into typed categories (correct, partially correct, a specific misconception, off-topic) with a confidence for each. Claude writes personal feedback only where it is needed and sends the low-confidence cases to the teacher, whose time goes to the students who need it.
Course evaluations and forum posts. Every semester, a university or school district collects thousands of free-text comments that nobody has time to read. Jev classifies each one by topic (workload, clarity, the professor, the platform) and by sentiment, and flags forum posts from students who sound lost or frustrated. Claude writes the summary for each department and the teacher gets the list of students to talk to this week.
Healthcare
These are ideas for research and decision support, never replacements for clinicians. Synthetic personas are not patients, and regulation, privacy and human oversight all apply.
Public health messages tested before launch. Model a population across age, health literacy, trust in institutions, faith and language. Claude drafts versions of a screening campaign, and Jev judges, per persona, whether it would act and what objection it would raise. Voxari's "would you leave" becomes "would you get screened", and you can find the message that works for the skeptical province, not only for the average.
Triage and routing support. Intake notes and patient portal messages classified into urgency, likely department and "needs human review" in about 300 milliseconds, with low-confidence cases always going to a nurse. Claude writes the summary for the clinician.
Patient feedback and incident reports. Surveys, complaints and incident reports classified by topic, sentiment and whether they are safety-related, as they arrive. Patterns show up in days instead of at the end of the quarter, and Claude writes the report for the quality team.
Clinical trial screening. Eligibility criteria become typed yes/no questions, Jev checks thousands of de-identified records against them at almost no cost, and Claude explains each match to the coordinator. Recruitment is one of the slowest parts of research.
Checking the writer. When Claude drafts a discharge summary or a patient letter, Jev checks each claim against the chart with a yes/no question and a probability. Anything under the threshold goes back to a human before it leaves the building. The same trick works for citations in a literature review. This one flips the roles: the judge is grading the thinker.
Manufacturing
A synthetic workforce for change management. Before changing shift patterns or bringing in automation, model the workforce by role, seniority, shift, site and trust in management, and see who resists and why. Voxari's "trust in the council" becomes "trust in management". Then go talk to the real people with better questions, before the all-hands where everybody hears about it for the first time.
Maintenance logs and inspection reports. Plants sit on years of free-text technician notes. Jev classifies each one by failure mode, severity and safety relevance, with a confidence. Those labels become columns in a table, and that table feeds a normal predictive model for the next failure. Claude reads the clusters and writes the root-cause analysis for the plant manager.
Warranty claims and customer complaints. Every claim classified by defect type and by how angry the customer is, then matched to the production batch. When one line starts generating the same complaint, somebody knows before the recall.
Shift notes and near-miss reports. Operators write short notes at the end of every shift, and most near misses live only there. Jev classifies each one by hazard type and severity as it is typed, so the safety team sees a pattern in the week it starts instead of after the accident. Claude writes the weekly safety summary.
One idea cuts across all of these. Claude plans, and Jev works as the cheap judge inside the loop, scoring thousands of items and calling Claude only when the judgment is uncertain. Knowing when the model is not sure is what lets you send the hard cases to a human.
Voxari are not real people
The README of Voxari ends with this line: "Synthetic opinion is not opinion."
The Voxari are invented. Synthetic populations are for rehearsing and for finding better questions. They never replace asking real people, and that goes double for patients, students and workers.
Jev answers confidently even when a persona has no real basis to decide. That is why Voxari forces an "undecided" option on every question and shows the confidence next to every result.
Personas carry the assumptions of whoever wrote them. Biased personas give biased answers. In healthcare and education, that is a real risk and it needs consideration.
The questions and the state sent to Jev have to be designed carefully. That design work is where Claude shines.
In sensitive domains, humans stay in the loop, and the low-confidence cases go to them.
Neither tool does the job alone. Claude can't answer as 5,000 people in nine seconds for ten cents, and Jev can't design a world or build a web app.
A Quick Demo
Are you curious about Voxari? I recommend going to voxari.org and give it a try. You can also check my quick demo on the video below.
Go ask them something
The Voxari are waiting at voxari.org. Send them a message and watch the Lattice light up. Read the Chronicle, look up a citizen by name, and hover over the speech bubbles until you learn a few words of Voxari. If your proposal wins by enough, your words become part of their history.
We keep talking about "AI" as if it were one thing. Building Voxari made it obvious to me that the interesting future is different kinds of AI working together: a thinker that plans and writes, a judge that decides instantly and at scale, and a human who chooses what is worth asking.
The Voxari have a line about that. The creator is said to still be listening. So go ahead, ask them something. The thinker, the judge and I will be listening.




Comments