Multipass AI
Multipass AI started as a Stanford exploration into a core challenge: Large language models can be fluent without being reliably true. The product thesis reframes trust as something earned through transparency, alignment, and visible dissent.
Product arc
Define the problem through AI product coursework and trust-centered research.
How might we build a tool that empowers users to trust AI answers through consensus?
Separate consensus workflows from creative workflows based on user intent.
Design long-term retrieval, semantic recall, and expiration models.
Model sustainable pricing, infrastructure cost, and long-term product evolution.
AI Hallucination is a problem
The original premise was straightforward: if models are probabilistic, the experience should help users inspect alignment, disagreement, and provenance instead of obscuring them behind a single polished answer.
That insight opened a larger design space around consensus, creativity, trust calibration, and when divergence becomes a feature rather than a flaw.
Unit economics
Simplified from the working spreadsheet: conversion, API cost, Stripe fees, infrastructure, and tier pricing were modeled together so the product could stay trustworthy and financially sustainable.
Pricing model
The business and customer problems both represent design problems. Example: Do customers understand token costs, limits, context windows that virtually all competitors base their charges on? Overwhelmingly, the answer is no.
Multipass charges by the question - not meaningless tokens.
Memory architecture
Classify prompts up front so memory budget and search behavior can adapt to the task.
Combine semantic and lexical signals, embeddings, reranking, and temporal expiration.
Surface the right memory at the right moment rather than loading everything equally.
Preserve freshness, reduce noise, and create a cleaner long-term interaction loop.
Universal cache
One of the most compelling future-state concepts is a universal cache: a consensus-vetted repository that captures reusable knowledge instead of letting high-value work disappear inside isolated chat sessions.
This module can become one of the signature visual moments on the live site, using provenance labels and trust states to show how knowledge compounds over time.
What we've learned
Across live comparisons, disagreement persisted. That is the point: users need visibility into where models align, diverge, and when fact-checking earns trust.
Observed disagreement ranged from 10.6% to 21.2% across the evaluation set, reinforcing the need for consensus-aware UX rather than one polished answer.
New since launch
Supporting dissent
Next steps
The direction forward is not only more features. It is a more durable system where verified answers, reusable artifacts, and trusted context compound over time.