Case Study Multipass AI

Multipass AI

Designing trust in AI through consensus, memory, and transparent reasoning.

Multipass AI started as a Stanford exploration into a core challenge: Large language models can be fluent without being reliably true. The product thesis reframes trust as something earned through transparency, alignment, and visible dissent.

Tool WeWeb Front End
Tech Supabase Back End
Workflow Cursor AI
Workflow Claude Code
System Hybrid RAG Memory
Focus AI Hallucination Reduction
0
Users on the platform since Q4 2025 launch (as of Q1 2026)
0
Major AI models included in consensus, plus a 6th for synthesis and agreement calculation
3 tiers
Plus, Premium, and Pro business model framing after a 10 query / 30 day free trial
Live thesis
An actively growing trust-first product that now includes fast modes and image generation

Product arc

From research to platform.

01

Stanford incubator

Define the problem through AI product coursework and trust-centered research.

02

Discovery & Definition

How might we build a tool that empowers users to trust AI answers through consensus?

03

Model behavior

Separate consensus workflows from creative workflows based on user intent.

04

Memory

Design long-term retrieval, semantic recall, and expiration models.

05

Scale path

Model sustainable pricing, infrastructure cost, and long-term product evolution.

AI Hallucination is a problem

Most AIs hide uncertainty. Multipass empowers users to verify answers through consensus.

The original premise was straightforward: if models are probabilistic, the experience should help users inspect alignment, disagreement, and provenance instead of obscuring them behind a single polished answer.

That insight opened a larger design space around consensus, creativity, trust calibration, and when divergence becomes a feature rather than a flaw.

Unit economics

Pricing and profitability were modeled from real query cost, not intuition.

Modeled tier
Free
Plus
Premium
Price / mo
$0
$10
$100
Query cap
10
100
1,000
Est. cost / user
$1.32
$1.65
$4.90
Gross margin
Acquisition
83.5%
95.1%
Scale read 70%+ net margin in mature modeled scenarios
Cost logic Blended cost API, hosting, and ops factored into pricing

Simplified from the working spreadsheet: conversion, API cost, Stripe fees, infrastructure, and tier pricing were modeled together so the product could stay trustworthy and financially sustainable.

Pricing model

Three tiers calibrated to usage volume, cost structure, and trust-critical workflows.

  • Free tier for lightweight trial, awareness, and initial trust building.
  • Plus tier at $10 / month Up to 100 queries with modeled 83.5% gross margin.
  • Premium tier at $100 / month Up to 1,000 queries with modeled 95.1% gross margin.
  • Pro tier at $1000 / month Up to 10,000 queries with modeled 96.2% gross margin.
  • Scale scenarios were benchmarked against infrastructure cost, payment fees, and market norms.

The business and customer problems both represent design problems. Example: Do customers understand token costs, limits, context windows that virtually all competitors base their charges on? Overwhelmingly, the answer is no.

 

Multipass charges by the question - not meaningless tokens.

Memory architecture

A layered system for relevance, freshness, and retrieval quality.


Routing

Classify prompts up front so memory budget and search behavior can adapt to the task.

Storage

Combine semantic and lexical signals, embeddings, reranking, and temporal expiration.

Recall

Surface the right memory at the right moment rather than loading everything equally.

Quality

Preserve freshness, reduce noise, and create a cleaner long-term interaction loop.

Universal cache

A shared knowledge layer for reusable facts, solved problems, and human-AI collaboration.

One of the most compelling future-state concepts is a universal cache: a consensus-vetted repository that captures reusable knowledge instead of letting high-value work disappear inside isolated chat sessions.

 

This module can become one of the signature visual moments on the live site, using provenance labels and trust states to show how knowledge compounds over time.

What we've learned

Agreement patterns make consensus value tangible.

Across live comparisons, disagreement persisted. That is the point: users need visibility into where models align, diverge, and when fact-checking earns trust.

Claude
78.8%
Gemini
80.6%
Llama
83.0%
ChatGPT
88.1%
Grok
89.4%

Observed disagreement ranged from 10.6% to 21.2% across the evaluation set, reinforcing the need for consensus-aware UX rather than one polished answer.

New

New since launch

  • Single AI mode across ChatGPT, Claude, Gemini, Grok, Llama, and Fast Mode.
  • Fact check button for single-AI mode that runs all five major models in parallel.
  • Advanced image editing and generation workflows.
  • Document, spreadsheet, and presentation editing and generation.

Supporting dissent

Next steps

Transform validated trust patterns into durable workflows and daily utility.

  • Design an agent that auto-populates and maintains high-confidence cache answers.
  • Add project organization so users can separate research, build streams, and long-running work.
  • Support persistent artifacts for code, novels, documents, and other living outputs.
  • Launch a dedicated fact-checked news pulse powered by consensus verification.

The direction forward is not only more features. It is a more durable system where verified answers, reusable artifacts, and trusted context compound over time.