System Architecture · Module Reference · Vision Paper · 2025

Commonsent:
A Modular Architecture for
Collective Sense-Making

Eight interlocking modules that together form a second perceptual interface, a deliberately designed prosthetic for navigating a world our evolved cognition can no longer read unaided.

Document typeReference architecture & use-case catalogue
Modules8 functional · 1 governance layer
Reading timeApprox. 45 minutes
Figures13 diagrams & graphs
Executive Summary

The human perceptual system is an evolved interface, optimized over hundreds of thousands of years for a world of small groups, visible threats, and short time horizons. We no longer live in that world. The environment we have built, networked, statistical, adversarial, operating at planetary scale and machine speed, produces risks and opportunities that our inherited cognition cannot render. The gap between what we can perceive and what survival now requires is the central problem this architecture addresses.

Commonsent is proposed as a second interface layer: a collective perceptual prosthetic assembled from eight interlocking modules. Each module compensates for a specific failure mode of the evolved headset. Together they form a system that can verify the origins of information, synthesize distributed signals into navigable maps, model trust at scale, detect emerging systemic risk, coordinate deliberation beyond the limits of face-to-face groups, render complexity directly into the perceptual field, mediate the individual's relationship with the commons, and govern itself without becoming a new instrument of capture.

This paper specifies each module, its function, its mechanics, and the concrete situations in which it earns its place, and shows how the modules compose into workflows greater than the sum of their parts. The architecture described here is a design proposal. The accompanying graphs are illustrative projections that express the intended behaviour of the system, not measurements of a deployed one.

Contents
Part I, The Problem and the Design
  1. 01The Interface Problem, in Brief
  2. 02Five Design Principles
  3. 03System Overview: How the Modules Fit Together
Part II, The Modules
  1. M1Provenance & Attribution
  2. M2Synthesis & Sense-Making Engine
  3. M3Trust Graph & Reputation
  4. M4Signal Detection & Early Warning
  5. M5Collective Deliberation & Coordination
  6. M6Perceptual Rendering & Spatial Interface
  7. M7Personal Cognitive Companion
  8. M8Commons Governance & Stewardship
Part III, Integration and Outlook
  1. 04Cross-Module Workflows: Three Scenarios
  2. 05Adoption Pathway
  3. 06Risks, Failure Modes, and Limits
  4. 07Conclusion
PART I
The Problem and the Design
Why a second interface, and the principles that shape it
01 · Foundation

The Interface Problem, in Brief

Perception is not a window onto reality. It is an interface evolved for fitness, and an interface can become obsolete when the environment it was built for disappears.

The cognitive scientist Donald Hoffman has argued, across three decades of work in evolutionary game theory and the philosophy of perception, that natural selection does not favour organisms that perceive reality accurately. It favours organisms that perceive it usefully. The result is what he calls a species-specific interface: a compressed, actionable rendering of a deeper structure, in the way that the icons on a computer desktop are useful stand-ins for circuitry no user needs to understand. Three-dimensional space, solid objects, faces, the flow of time, these are our icons. They are real as icons. They are not the underlying machine.

This interface was calibrated for a particular world: bands of roughly a hundred and fifty people, threats you could see and run from, trust signals carried by face and voice, causal chains short enough to follow from beginning to end, and time horizons measured in days and seasons. Within that world the interface is breathtakingly good.

We do not live in that world. We built a different one and walked into it wearing the old headset. The new environment is dominated by exactly the things the evolved interface cannot render: statistical and invisible risks that arrive before any symptom appears; actors who are not physically present and may not be human; trust signals systematically forged at industrial scale; causal chains that loop across continents and decades. The headset is not malfunctioning. It is running flawlessly, and rendering the wrong things with perfect fidelity. The threat does not look like a threat. The fabrication does not feel false. The slow-building collapse does not register until it arrives as a shock.

When an interface fails this way, reasoning harder inside it does not help; you cannot think your way past a perceptual limit using the very faculty that is limited. What is required is a new interface layer, a prosthetic that renders, in tractable form, the parts of reality the biology leaves invisible. This is what tools at their most consequential have always been. The telescope extended sight. Writing extended memory. Commonsent is proposed as the same kind of extension, aimed at the faculty whose failure is now most dangerous: our collective capacity to perceive shared reality and act on it together.

One clarification guards against a common misreading. To say the interface is a rendering rather than a window is not to say that anything goes, or that all renderings are equally good. The whole force of the evolutionary argument runs the other way: interfaces are ruthlessly graded by the environment, and an interface that misleads its bearer about fitness-relevant structure gets that bearer killed. The savanna interface was not arbitrary; it was exquisitely fitted to its world. The problem is not that perception is constructed — all perception is — but that our particular construction was fitted to a world we no longer occupy. The remedy is therefore not to abandon interfaces in pursuit of some impossible unmediated vision, but to build a better-fitted one. Commonsent makes no claim to show reality as it is. It claims only to render the parts of our actual environment that the inherited interface leaves dark, and to render them well enough that the people wearing it survive choices the old interface would have fumbled.

This reframing also dissolves a debate that otherwise traps these discussions. Asked whether some new technology will "save us" or "doom us," people reach for sweeping verdicts about technology in general. The interface lens asks a sharper, answerable question instead: does this particular construction render the relevant environment more or less faithfully than what we have now? A tool that sharpens collective perception of real structure is good in the specific, measurable sense that its bearers make better-fitted choices; a tool that renders engaging fictions is bad in the same specific sense, however pleasurable. The point of an explicitly designed interface is that, unlike the evolved one, its fitness can be examined, contested, and corrected on purpose. That is the whole wager of building the second headset deliberately rather than waiting another ten thousand generations for the first one to catch up — time the present environment does not give us.

02 · Design

Five Design Principles

If the diagnosis is an interface problem, then the cure is not "more information." A failing interface delivered more raw data fails faster. The architecture is governed instead by five principles, each a direct response to a way that naïve information systems make the problem worse.

Principle 1 · Render, don't deliver

The system's job is not to hand people facts but to change the interface through which they perceive. A figure buried in a report changes little; the same figure rendered as a vivid, spatially located feature of the perceptual field changes behaviour. Every module is evaluated by whether it improves the quality of perception, not the quantity of data moved.

Principle 2 · Provenance before content

In an environment where fabrication is cheap and ubiquitous, the origin of a claim is more load-bearing than the claim itself. The architecture treats provenance as a first-class property of every piece of information, computed and carried before content is ever synthesized or shown.

Principle 3 · Distributed trust, never central authority

A system that becomes the single arbiter of truth becomes the single point of capture and the most valuable target an adversary could wish for. Trust in Commonsent is modelled as a contextual, distributed, auditable graph, not a verdict issued from a centre.

Principle 4 · Legibility of its own workings

A prosthetic that the user cannot inspect is indistinguishable from a manipulation. Every weighting, every synthesis, every alert must be traceable to the reasons that produced it. The system explains itself as a condition of being trusted.

Principle 5 · Compose, don't centralize

The modules are designed to interoperate through shared interfaces rather than to fuse into a monolith. This keeps the system adaptable, allows components to be replaced or governed independently, and prevents any one module from quietly accumulating power over the others.

These principles are constraints, not features. Each one closes a door that a careless version of the same system would leave open, the door to manipulation, capture, opacity, or the simple amplification of noise.

The principles as a system, not a list

It is worth seeing how tightly these five interlock, because their power is in the combination rather than in any one alone. Render-don't-deliver (1) would be a manipulation engine without legibility (4) to keep its renderings honest. Provenance-before-content (2) would be inert without distributed trust (3) to interpret what the provenance means. Distributed trust (3) would fragment into incoherence without the synthesis that render-don't-deliver demands. And compose-don't-centralize (5) is what keeps the other four from quietly fusing into a monolith in which a single team's judgment silently governs all of them. Remove any one principle and a failure mode the others were holding shut swings open. This is why the architecture treats them as binding simultaneously rather than as a menu from which a builder might pick the convenient ones, a system that honoured four of the five would not be four-fifths as safe; it would be unsafe along exactly the dimension it dropped.

The deeper point is that all five descend from a single source: the recognition that a perceptual prosthetic for a society is an instrument of power over what that society perceives. Each principle is a way of refusing to let that power concentrate, hide, or act without accountability. They are, in effect, the constitutional limits of the system — written first, precisely because everything built afterward will be tempted to exceed them.

03 · Architecture

Where this sits in the series. This document specifies the collective sense-making architecture: the perceptual interface layer of Commonsent. The economic and coordination layer, participant-owned treasuries, antitrust demand routing, and net-neutral transactional investment, is a separate layer of the same network. The eight modules here map onto the rest of the series. The perceptual interface in M6 is the subject of The Misfit Headset. Modules M1 to M5 are the technical components that implement the civic-intelligence division published as Commonsent Signal, where M4 signal detection is one component of that division rather than the division itself, and the decaying, topic-specific reputation in M3 is the Topic Credit Ladder described there. The deliberation in M5 follows the Commonsent Deliberation Protocol. What all eight modules ultimately optimize for is set out in The Objective Function.

System Overview: How the Modules Fit Together

The eight modules organize into four bands. At the base sits the foundation band, where information enters the commons and acquires verifiable identity: the Provenance & Attribution module (M1) and the Trust Graph (M3). Above it, the cognition band turns verified signal into understanding: the Synthesis Engine (M2) and Signal Detection (M4). The interaction band connects understanding to human action: Collective Deliberation (M5), Perceptual Rendering (M6), and the Personal Cognitive Companion (M7). Wrapping all of them is the stewardship band, Commons Governance (M8), which is not a layer the data flows through but a frame that holds the whole system accountable.

Information moves upward through the bands, gaining structure at each step, while governance presses inward on every module at once. The figure below shows the full stack.

Commonsent master architecture Four horizontal bands containing eight modules, with information flowing upward and governance wrapping all modules. M8 · COMMONS GOVERNANCE INTERACTION BAND M5 · Deliberation Coordinate at scale M6 · Rendering Spatial overlay / XR M7 · Companion Personal interface COGNITION BAND M2 · Synthesis & Sense-Making Engine Builds the collective epistemic map M4 · Signal Detection Surfaces emerging systemic risk FOUNDATION BAND M1 · Provenance & Attribution Verifiable origin for every claim M3 · Trust Graph Distributed, contextual reputation INPUTS Documents · media Sensors · data feeds Human testimony Institutions · AI signal gains structure ↑ to human perception & coordinated action
Figure 1. The Commonsent stack. Raw inputs enter at the bottom and rise through the foundation, cognition, and interaction bands, acquiring verifiable identity, then structure, then a form fit for human perception. The Commons Governance module (M8) is drawn as a dashed frame because it does not sit in the data path, it constrains every module simultaneously.

Two properties of this arrangement are worth stating plainly. First, nothing reaches a human eye until it has passed through the foundation band; there is no path by which unverified content can be rendered as if it were established. Second, governance is structurally inescapable, it is not a module you can route around, because it is the frame, not a step.

It is also worth being explicit about what the band structure buys in terms of resilience. Because the modules communicate through shared interfaces rather than shared internals, a failure or compromise in one need not propagate to the others: a degraded Synthesis Engine produces worse maps but cannot forge provenance; a captured rendering layer can mislead about presentation but cannot rewrite the trust graph beneath it. The architecture is, in this sense, designed to fail in parts rather than as a whole — a property that matters enormously for something meant to mediate perception at scale, where a single catastrophic failure mode would be intolerable. Decomposition into modules is not merely an engineering convenience for building the system; it is a safety property of the running one.

The remainder of this paper takes each module in turn. For each, the paper states what the module does, how it works, why it takes the form it does rather than a simpler-looking alternative, and the concrete situations — from the civilizational to the everyday — in which it earns its place. The use cases are deliberately drawn from different domains for each module, to show that these are general-purpose components of perception rather than point solutions to a single problem.

PART II
The Modules
Eight components, their mechanics, and where each earns its place
M1

Provenance & Attribution

Foundation band · trust origin layer

Before anyone asks whether a claim is true, the system asks where it came from, and makes that answer impossible to forge.

The evolved headset assigns credibility through face, voice, and physical presence. Those signals once correlated reliably with the actual source of information; an adversary could not easily wear another person's face. That correlation has collapsed. Synthetic media, large-scale account fabrication, and the industrial production of plausible text have severed the link between the trust cue and the trustworthy origin. The Provenance & Attribution module exists to rebuild that link on a foundation the biology never had: cryptographic verifiability.

How it works

Every artifact that enters the commons, a document, an image, a measurement, a statement, is content-addressed and bound to a signed attestation describing its origin: who or what produced it, when, and on the basis of what prior material. As the artifact is edited, quoted, summarized, or recombined, each transformation appends to a tamper-evident chain of custody. The chain does not assert that the content is correct; it asserts, verifiably, what the content is and where it has been. A reader confronting a claim can therefore see whether it originates in primary observation, in a chain of citations terminating in a credible source, or in a loop of mutually citing fabrications, a structure the unaided interface renders identically in all three cases.

Crucially, provenance is computed and carried before content is synthesized (Module 2) or shown (Module 6). This ordering is the practical expression of design Principle 2. A synthesis built on unverifiable inputs is flagged as such at every downstream step, rather than laundering its uncertain origins into apparent authority.

Provenance chain of custody A chain showing an artifact's origin, transformations, and a forked branch flagged as unverifiable. SRC signed Primary source EDIT Transformation CITE Quoted in report Verified claim chain intact ? Broken / circular origin flagged before it reaches a reader
Figure 2. A chain of custody. The unaided interface renders a sourced claim and a fabricated one identically. M1 makes the difference structurally visible: an intact, signed chain terminating in a primary source resolves to a verified claim, while a circular or broken origin is flagged at the point it tries to enter circulation.

Why provenance, and not a truth label

The intuitive alternative is to label claims true or false. That approach fails for two reasons the interface lens makes obvious. First, it requires an arbiter, and an arbiter of truth is precisely the central authority that design Principle 3 forbids, a single point of capture and the richest target an adversary could want. Second, truth is frequently unsettled at the moment a claim circulates; a binary label forces a verdict the evidence does not yet support, and either freezes a premature answer or discredits the labeller when it proves wrong. Provenance sidesteps both traps. It makes no claim about correctness, only about origin, which is a verifiable fact rather than a contested judgment. A reader can then apply their own standards, supported by trust weighting (M3) and synthesis (M2), to material whose pedigree is no longer in doubt. The module raises the floor of accountability without appointing anyone the ceiling.

Where it earns its place

Newsroom
Verifying a leak

A journalist receives a leaked document. Before publishing, they trace its chain of custody, confirm it terminates in a genuine primary source rather than a plausible forgery, and attach that provenance to the published story so readers can check it too.

Science
Living meta-analysis

A meta-analysis automatically traces every input study's provenance and flags any that have since been retracted or corrected, so the synthesis updates the moment its foundations shift rather than years later.

Crisis response
Field report triage

During a disaster, responders are flooded with reports. The module distinguishes verified field observations from unsourced rumour in real time, letting limited attention flow to signals that can actually be acted on.

Media integrity
Synthetic-media disclosure

An image carries an attestation that it was machine-generated. That fact travels with the image and is surfaced at the moment of perception, rather than depending on a viewer's ability to spot artifacts the technology is rapidly eliminating.

M2

Synthesis & Sense-Making Engine

Cognition band · the collective epistemic map

No individual mind, and no single institution, can hold the full state of a complex situation. This module holds it for them, and hands back a map they can read.

The evolved interface is built for situations a single nervous system can encompass. The defining problems of the present escape that bound: a pandemic, a financial system, a contested public question debated across millions of fragmentary sources. The Synthesis Engine ingests provenance-tagged inputs and constructs a navigable representation of what is known, what is contested, and what is genuinely unknown, the collective epistemic map.

How it works

The engine models a domain not as a pile of documents but as a structured field of claims and the relationships between them: which claims support which, which contradict, which are independent, and how strongly each is anchored in verified provenance. Large language models do the heavy lifting of reading, extracting, and relating, but the output is not a confident summary that hides its own seams. It is a map that preserves disagreement, marks confidence honestly, and lets a reader zoom from a one-line overview down to the specific evidence, and the specific gaps, underneath any point on it.

This honesty about uncertainty is a deliberate inversion of how most systems behave. A search engine or chatbot is rewarded for producing a single fluent answer; the seams of disagreement are sanded away. The Synthesis Engine is rewarded for the opposite: rendering the shape of a question accurately, including the places where the answer is genuinely unsettled, so that a decision-maker is not misled by manufactured consensus.

Synthesis funnel Many scattered sources converging through a synthesis engine into a structured epistemic map with known, contested, and unknown regions. scattered sources Synthesis engine AI + provenance EPISTEMIC MAP Known & well-supported Contested, open disagreement Unknown, gaps marked honestly
Figure 3. The Synthesis Engine converges scattered, provenance-tagged sources into a structured epistemic map. Unlike a single fluent answer, the map preserves the distinction between what is settled, what is genuinely contested, and what is unknown, the three regions a decision-maker most needs to tell apart.

Why a map, and not an answer

Every commercial incentive pushes toward the single fluent answer: it is what users say they want, what fits a chat box, and what demos well. The interface lens exposes the cost. A confident answer that conceals genuine disagreement is not a better interface to reality; it is a more convincing one to a fabricated reality, which is strictly worse. When the underlying question is contested, the honest rendering of that question is the contest, and a decision-maker who is shown false consensus makes worse decisions than one shown true uncertainty, because they cannot price the risk they cannot see. The Synthesis Engine therefore optimizes for fidelity to the shape of a question, not for the comfort of resolution. Where an answer exists and is well-supported, the map says so plainly; where it does not, the map refuses to manufacture one. This is less satisfying and far more useful.

Where it earns its place

Policy
Live evidence map

A legislature considering a bill receives a continuously updated map of the evidence for and against it, with each strand weighted by provenance and confidence, and the points of real disagreement made explicit rather than averaged away.

Medicine
State of the evidence

A clinician sees the synthesized state of evidence on a treatment, updated as new trials land and old ones are revised, instead of relying on a guideline that may be years behind the literature.

Investigation
Mapping a network

Investigative journalists map a sprawling corruption network across tens of thousands of documents, with the engine surfacing the connections, contradictions, and missing links no single reporter could hold in mind.

Public discourse
De-fogging a debate

On a polarized public question, the map shows readers what the disagreement is actually about, separating empirical disputes from value differences, which the unaided interface fuses into a single undifferentiated fight.

M3

Trust Graph & Reputation

Foundation band · distributed credibility

Not "who is trustworthy" decided from a centre, but "trusted by whom, for what, on what record", modelled openly enough that anyone can see why.

Provenance (M1) tells you where a claim came from. The Trust Graph tells you how much weight that origin has earned, in a specific domain, based on a track record anyone can inspect. It is the module most exposed to the temptation that design Principle 3 forbids: the temptation to become an authority. A system that issued a single trust score per source would be both brittle and dangerous, brittle because trust is contextual, and dangerous because a single score is a single lever an adversary or owner could pull. The Trust Graph instead models trust the way it actually exists: as a web of contextual, directional relationships.

How it works

Reputation in the graph is domain-specific. A virologist's testimony carries weight on viral transmission and little on monetary policy; the graph keeps those separate rather than collapsing them into a generic authority score. Trust propagates along edges, if sources you have reason to trust consistently rely on a third source, that flows through the graph, but propagation decays with distance and is always traceable to its grounds. And every weighting is accompanied by its reasons: a reader can ask why a source is weighted as it is and receive an answer in terms of track record, domain, and the chain of relationships that produced the figure, never an opaque verdict.

Trust also decays. A reputation earned a decade ago in a fast-moving field is not the same asset as one demonstrated last month, and a source that was reliable until it was captured or corrupted should lose standing as the evidence of that shift accumulates. The graph models this decay explicitly rather than treating reputation as a permanent endowment.

Trust graph and reputation decay Left: a network of contextual trust edges between sources. Right: a curve showing reputation decaying over time without renewal. CONTEXTUAL TRUST EDGES A B C D Edge thickness = trust weight in a given domain. No central score; trust is read off the web. Every edge carries its reasons. REPUTATION DECAY WITHOUT RENEWAL high low time → renewed by track record left to decay
Figure 4. Left, trust is read off a web of contextual, weighted edges rather than issued as a single score, keeping the system from becoming an authority. Right, illustrative reputation dynamics: standing decays without renewal (rose) but is sustained by a continuing track record (green dashed), so a reputation cannot become a permanent, unearned endowment. Illustrative; not measured data.

Why a graph, and not a score

A single trust score per source is seductive because it is simple to display and simple to act on. It is also a weapon. Whoever sets the score governs perception, and any source can be destroyed or elevated by moving one number. The graph refuses this simplicity deliberately. By keeping trust contextual, domain-specific, directional, and accompanied by its grounds, it denies any actor a single dial to turn, and it matches how trust actually behaves: a person you would believe completely about their own field you would not believe at all about another. The cost is that trust in Commonsent can never be reduced to a tidy badge. That cost is the feature. A system whose trust signals could be compressed into one number could be captured by capturing that number.

Where it earns its place

Expertise
Signal vs credential

The graph surfaces genuine, demonstrated domain expertise, including from people without conventional credentials, while down-weighting credentialed voices speaking outside their actual track record.

Communities
Earned moderation weight

An online community weights contributions by demonstrated reliability in that specific context, reducing the leverage of brigading and sock-puppetry without installing a censor.

Bridging
Crossing epistemic bubbles

The graph identifies sources trusted across otherwise separated communities, providing rare bridges that can carry credible information between groups that share no common authority.

Integrity
Detecting capture

When a once-reliable institution is captured or begins to drift, the decay dynamics register the change as its outputs stop matching its record, rather than letting an old reputation shield new failures.

M4

Signal Detection & Early Warning

Cognition band · surfacing the invisible

The most dangerous threats in the modern environment share one feature: they are imperceptible to the evolved headset until they are too large to stop.

A predator on the savanna triggers an immediate, vivid perceptual alarm. A pandemic in its first weeks, a financial contagion building in the links between institutions, a slow ecological threshold approaching, none of these trip any biological alarm at all. They are statistical, distributed, and silent. By the time they cross into the range the evolved interface can perceive, the cheap interventions are gone. The Signal Detection module is built to perceive these threats early, on humanity's behalf, in the window where action is still inexpensive.

How it works

The module runs anomaly detection across many provenance-tagged data streams at once, looking not for any single dramatic event but for weak, correlated signals, patterns that are individually unremarkable but jointly significant. It amplifies these weak signals to the threshold of attention while explicitly modelling its own false-positive rate, because an early-warning system that cries wolf is quickly and rightly ignored. When a pattern crosses a calibrated threshold, the module raises it into the cognition band with its full provenance and confidence attached, so that the people who receive the warning can see exactly what it rests on.

The figure below illustrates the core value proposition: detection ahead of the curve.

Early warning detection ahead of the curve A line graph showing a weak signal rising and crossing a detection threshold well before the threat becomes perceptible to conventional recognition. magnitude time → actual threat magnitude Commonsent detection threshold conventional recognition threshold detected recognized recovered intervention window
Figure 5. The value of early detection. A growing systemic threat (rose) stays below the threshold of conventional recognition (brown) for a long time. By fusing weak correlated signals, M4 crosses its detection threshold (green) much earlier. The bracket marks the recovered intervention window, the period in which cheap action is still possible. Illustrative schematic; axes are qualitative.

Why weak signals, and not loud alarms

The obvious way to build a warning system is to watch for big, obvious events. But by definition, a big obvious event has already crossed the threshold the evolved interface can perceive on its own, the warning arrives too late to be cheap. The value of M4 lives entirely in the pre-perceptible window, which means it must work with signals that are individually unremarkable and only jointly meaningful. This is also why false-positive discipline is not a refinement but a survival requirement: a system that amplifies weak signals will, if undisciplined, amplify noise into a stream of cried wolves, and a warning system that is ignored is worse than none, because it consumes the attention it was meant to direct. The hard engineering problem here is not sensitivity but calibrated sensitivity, surfacing the real pattern early without drowning it in plausible-looking nothing.

Where it earns its place

Epidemiology
Outbreak before recognition

Correlated weak signals, search patterns, clinical notes, pharmacy data, wastewater readings, are fused into an outbreak warning weeks before official case counts would surface it.

Finance
Contagion risk

The module maps building stress in the interconnections between institutions, flagging the kind of hidden coupling that turns a local failure into a systemic one before it cascades.

Information security
Coordinated inauthenticity

Patterns of synchronized, inauthentic behaviour across accounts and platforms are detected as a coordinated operation rather than read, by the unaided interface, as organic public sentiment.

Ecology
Threshold warning

Long-horizon environmental signals that the evolved time-sense cannot integrate, gradual drift toward an ecological tipping point, are rendered as a tracked, dated approach to a threshold.

M5

Collective Deliberation & Coordination

Interaction band · sense-making at scale

Human groups coordinate beautifully up to about a hundred and fifty people. Above that, the evolved machinery for finding agreement breaks down, exactly where the hardest problems now live.

The number is not arbitrary. The anthropologist Robin Dunbar's work suggests a cognitive ceiling on the number of stable relationships a person can maintain, and our deliberative instincts, reading the room, tracking who agrees with whom, sensing an emerging consensus, were tuned for groups beneath that ceiling. The questions that now most need collective resolution involve thousands, millions, or a whole polity. The unaided interface, asked to make sense of a debate at that scale, falls back on the crude heuristics it has: tribe, volume, and the loudest available story. The Collective Deliberation module replaces those heuristics with structure.

How it works

Rather than running deliberation as an undifferentiated flood of opinions, the module maps the space of views: it clusters positions, identifies where people actually agree (often far more than the surface suggests), and locates the precise cruxes, the specific points on which a genuine disagreement turns. It distinguishes, as Module 2 does for evidence, between empirical disputes that more information could resolve and value differences that it cannot. And it does this asynchronously and at scale, so that a deliberation among ten thousand people produces not ten thousand disconnected comments but a navigable map of the collective's actual structure of agreement and disagreement.

The effect is to make a large group's deliberative state perceptible, to render, for the first time, the shape of what a crowd thinks in a form a participant can read at a glance, the way a small group can read its own room.

Consensus and crux mapping Opinion clusters with a large shared-agreement region and a single identified crux point of genuine disagreement. shared agreement (often larger than it appears) Cluster A Cluster B The crux the one point the disagreement actually turns on 10,000 comments → one readable map empirical vs value disputes separated
Figure 6. Instead of presenting a deliberation as an undifferentiated flood, M5 maps it: most participants share a large region of agreement, two clusters diverge, and the genuine disagreement reduces to a single identifiable crux. Surfacing the crux is what makes large-scale deliberation tractable.

Why map the disagreement, and not tally votes

The reflexive tool for large-group decisions is the vote, or its softer cousin, the sentiment poll. Both compress a rich structure of views into a scalar, and in doing so destroy exactly the information a group needs to actually resolve anything. A 52 to 48 vote tells you a polity is split; it does not tell you on what, whether the split is empirical or moral, or whether the two sides agree on far more than the framing admits. M5 treats compression as the enemy. By mapping rather than tallying, it preserves the structure, the large shared region, the small genuine cruxes, that makes resolution possible. A group that can see it agrees on nine of ten points can negotiate the tenth; a group shown only the aggregate sees only an enemy. The module's job is to give a crowd the self-perception a small room has for free.

Where it earns its place

Democracy
Citizens' assembly at scale

Thousands deliberate a contested policy and the system shows, in real time, where they already agree and the few cruxes that remain, turning an unmanageable town hall into a structured collective decision.

Organizations
Distributed alignment

A globally distributed team aligns on strategy without endless meetings, as the module surfaces hidden agreement and pinpoints the genuine decisions that need a human call.

Conflict
Finding common ground

Between polarized groups, the map reveals the substantial common ground that tribal framing conceals, and isolates the real differences so they can be addressed honestly rather than performed.

Standards
Open protocol design

A large open community converges on a technical or governance standard, with the deliberation map preventing a vocal minority from being mistaken for consensus.

M6

Perceptual Rendering & Spatial Interface

Interaction band · the literal headset

Everything the other modules produce is worthless if it stays trapped on a screen the user must remember to consult. This module puts it where perception actually happens.

This is the most literal expression of the interface thesis. The whole framework rests on the claim that fitness is determined by the quality of the interface, not the quality of reasoning performed within it. A threat rendered as a vivid, spatially located feature of the perceptual field is responded to faster, more appropriately, and with less cognitive overhead than the same threat presented as an abstract figure on a separate display. The Perceptual Rendering module, including the actual wearable headset, exists to move the commons from information you look up to a layer on the world you inhabit.

How it works

The module takes the structured outputs of the cognition and interaction bands, verified provenance, epistemic maps, deliberative state, early warnings, and renders them as contextual overlays on the user's perceptual field through spatial computing. It is attention-aware: rather than flooding the field with annotation, it surfaces what is relevant to the user's current context and goals, and recedes otherwise. And it works in both directions of the interface thesis: it can render abstract, invisible structure (a financial coupling, a credibility weight, a deliberative crux) as something perceptually present, and it can locate information spatially on the physical world where that placement aids understanding.

Perceptual overlay concept A perceptual field with contextual overlays surfacing provenance, credibility, and warnings directly onto perceived objects. PERCEPTUAL FIELD, CONTEXTUAL OVERLAY A claim in view ✓ verified trust 0.82 An image in view ⚠ machine-generated A system in view early-warning: rising Attention-aware: surfaces what is relevant now, recedes otherwise. Invisible structure becomes perceptually present.
Figure 7. A conceptual perceptual overlay. Verified status, contextual trust weight, synthetic-media disclosure, and live early-warning state are surfaced onto the objects they pertain to, at the moment of perception, the difference between a fact you must remember to check and one you simply see.

Why the perceptual field, and not a better dashboard

One could object that all of this could live in a very good app, a screen you consult. The interface thesis says otherwise, and says it sharply: fitness is set by the quality of perception, not by the quality of information available somewhere to be looked up. A fact on a screen you must remember to check is, for most practical purposes, a fact you do not have at the moment of action. The whole evolutionary point of perception is that it operates without deliberate retrieval, you do not decide to see the predator. Moving the commons into the perceptual field is therefore not a convenience upgrade over a dashboard; it is a change in kind, from information that competes for attention to information that arrives as attention. The risk, of course, is a perceptual field so cluttered it degrades rather than aids perception, which is why attention-awareness is not optional polish but the core engineering constraint of the module.

Where it earns its place

Field science
Data on the world

A researcher in the field sees sensor data, historical records, and model predictions overlaid on the physical environment they are studying, integrating layers of information that would otherwise live in separate tools.

Decision rooms
Systemic state, spatially

Leaders in a crisis see the state of a complex system, supply chain, grid, epidemic, rendered spatially, letting a group perceive and discuss the same shared picture rather than competing dashboards.

Everyday
Provenance at the point of perception

Reading the news or scrolling a feed, a person sees provenance and credibility surfaced inline, so the trust judgment the biology can no longer make reliably is supported exactly where it is needed.

Education
Making abstraction visible

A complex abstract structure, a proof, a system, a historical causal web, is rendered as a navigable spatial object, using the same evolved spatial cognition that the unaided interface cannot apply to abstraction.

M7

Personal Cognitive Companion

Interaction band · the individual's interface to the commons

A commons that everyone must approach on its own terms serves no one well. The companion is the part of the system that knows you, and protects you, including from the system itself.

The other modules operate on shared, collective material. The Personal Cognitive Companion is the individual-facing agent that mediates between a particular person and the commons: it translates the collective epistemic map into terms that fit this person's context, knowledge, and goals, and it guards the one resource the modern environment most aggressively exploits, attention. Where the rest of the architecture asks "what is true, and how do we know," the companion asks "what does this person need to perceive right now, and how can it be rendered for them."

How it works

The companion learns a person's context, what they are working on, what they already understand, what they care about, and uses that to filter and frame the commons on their behalf. It scaffolds understanding of complex topics rather than dumping conclusions, raising the person's own capacity over time instead of fostering dependence. And, critically, its loyalty is structurally bound to the user, not to any party trying to capture their attention or belief: it is designed to protect against manipulation, including manipulation routed through Commonsent itself. This is the individual-scale embodiment of design Principles 1 and 4, rendering, and legibility, turned toward a single person's cognition.

Personal companion mediation loop A loop showing the companion mediating between the person and the commons, filtering inward and scaffolding outward. Person context · goals Companion loyal to user guards attention The commons M1, M6 outputs filter scaffold need / query retrieve The companion raises the person's own capacity rather than fostering dependence.
Figure 8. The companion's mediation loop. The person expresses a need; the companion queries and retrieves from the commons, then filters and scaffolds what returns into a form fit for this person, always loyal to the user, including against attempts to manipulate them through the system itself.

Why a loyal agent, and not a neutral tool

A tempting framing is that the companion should be neutral, a transparent pane of glass onto the commons. But neutrality, in an adversarial attention economy, is a fiction that benefits whoever is least neutral. Every other agent reaching for a person's attention is optimizing for something, engagement, persuasion, sale, and a companion that declines to optimize for the user simply cedes the field to those who optimize against them. So the companion is deliberately partisan in exactly one direction: toward its user's own considered goals, including against manipulation routed through Commonsent itself. This is also why it must scaffold rather than answer. An agent that maximizes a user's short-term satisfaction would foster the dependence that atrophies the very faculties the system exists to strengthen. Loyalty to the user's stated wants and loyalty to the user's actual flourishing can diverge, and the companion is specified to serve the latter.

Where it earns its place

Attention
A defended information diet

The companion helps a person navigate a hostile information environment without being captured by engagement-maximizing systems, acting as an advocate for their actual goals rather than their impulses.

Learning
Scaffolded understanding

Facing a complex topic, a person is guided to genuine understanding through their existing knowledge, rather than handed a conclusion that leaves them no more capable than before.

Decisions
High-stakes personal choices

For a major medical, financial, or life decision, the companion assembles the relevant verified evidence and the structure of the choice, in the person's terms, without substituting its judgment for theirs.

Accessibility
An interface for every capacity

The companion adapts the commons to a person's particular cognitive and sensory profile, making collective intelligence reachable regardless of how that individual best perceives and processes.

M8

Commons Governance & Stewardship

Stewardship band · the frame around everything

A system powerful enough to shape what a civilization perceives is powerful enough to be the most dangerous instrument ever built. Governance is not a feature of this architecture. It is its precondition.

Every prior module increases the system's influence over collective perception. That influence is the point, and the peril. A perceptual prosthetic for a whole society is, by construction, a mechanism for shaping what that society sees, trusts, and decides. If such a mechanism were owned, captured, or quietly tuned by any single actor, it would not solve the interface problem; it would become its most efficient possible exploitation. The Commons Governance module exists to make that capture structurally difficult, and to keep the system legitimate enough to be trusted with the role it plays.

How it works

Governance operates on three commitments. First, transparency of rules and workings: how the system weights, synthesizes, and surfaces is itself open to inspection, so that the prosthetic cannot become a black box that shapes perception for reasons no one can examine (design Principle 4). Second, distributed stewardship: control over the commons is spread across many parties with divergent interests, with no single point, owner, state, or faction, able to set its behaviour unilaterally (Principle 3). Third, exit and auditability: users can verify what the system is doing on their behalf, contest it, and leave with their data and relationships intact, so that legitimacy rests on genuine consent rather than lock-in. Decisions about the commons are themselves conducted in the commons, using the deliberation module, the system governs itself by the same means it offers everyone else.

Distributed stewardship structure The commons at the centre, governed by multiple distributed stewards with divergent interests, with transparency, audit, and exit as binding constraints. The commons no single owner Civil Public Users Domain BINDING CONSTRAINTS: transparency · audit · contestation · exit rights DISTRIBUTED STEWARDSHIP
Figure 9. Governance as distributed stewardship. Control over the commons is held by multiple parties with divergent interests, civil society, public institutions, users, and domain bodies, none able to act unilaterally, all bound by the same constraints of transparency, audit, contestation, and exit.

Why distributed stewardship, and not a benevolent owner

The cheapest governance story is a trustworthy owner, a wise foundation or a well-meaning company that promises to wield the system for good. The interface lens treats this as the most dangerous option available, precisely because it can be sincere. A benevolent owner is still a single point of capture, a single set of incentives that can drift, and a single target that an adversary, a market, or a state need only compromise once. Worse, benevolence is unfalsifiable from outside: a captured owner and an honest one make the same promises. Distributed stewardship replaces a promise of good behaviour with a structure that makes bad behaviour expensive and visible, many hands with divergent interests, none able to act alone, all bound to transparency and exit. It is messier, slower, and harder to build than a benevolent dictatorship of the commons. It is also the only version that a rational person should be willing to let mediate their perception of reality.

Where it earns its place

Capture resistance
No single lever

Because no one party can set the system's behaviour, an actor seeking to bend collective perception finds no single lever to pull, the cost and visibility of capture are raised by design.

Accountability
Governing in the open

Decisions about how the commons works are made through its own deliberation module, in public view, so the system's evolution is itself subject to the scrutiny it provides to everything else.

Legitimacy
Consent, not lock-in

Users can audit what the system does on their behalf and leave with their data and relationships intact, so participation rests on ongoing consent rather than captured dependence.

Adaptation
Legitimate evolution

As the environment changes, the rules of the commons can be revised through a process the community recognizes as legitimate, letting the prosthetic adapt without any single steward seizing the chance to reshape it.

PART III
Integration and Outlook
How the modules compose, and what it would take to build them
04 · Integration

Cross-Module Workflows: Three Scenarios

The modules are valuable individually, but the architecture's purpose is realized when they compose. A signal means little without provenance behind it; a deliberation is hollow without a shared evidence map beneath it; an overlay is dangerous without governance around it. The figure below shows which modules feed which, the connective tissue of the system, and the scenarios that follow trace concrete paths through it.

Module interaction matrix An 8 by 8 grid showing which modules feed into which others, with governance touching all. FEEDS INTO → FROM ↓ M1 M2 M3 M4 M5 M6 M7 M8 M1 M2 M3 M4 M5 M6 M7 M8 strong feed supporting feed governance reach
Figure 10. The module interaction matrix. Each filled cell shows that the row module feeds the column module. Provenance (M1) feeds everything that follows; the Synthesis Engine (M2) feeds the entire interaction band; and Governance (M8) reaches every functional module, the purple row across the bottom, without sitting in the data path.

Scenario A, An emerging outbreak

A novel pathogen begins spreading in a region. Signal Detection (M4), fusing weak signals across clinical, pharmacy, and wastewater streams, each carrying provenance (M1), crosses its detection threshold weeks before official case counts would. The Synthesis Engine (M2) assembles the fast-moving evidence into a live map, marking honestly what is known and what is still unknown, with each strand weighted by the Trust Graph (M3). Public-health bodies and affected communities use Deliberation (M5) to coordinate a response at a scale no single agency could convene, and the evolving situation is rendered to responders through Perceptual Rendering (M6) as a shared spatial picture. Individuals receive context-appropriate guidance through their Companion (M7), and the whole response runs under Governance (M8) that keeps any one actor from distorting the shared picture for advantage. The recovered intervention window of Figure 5 becomes weeks of lead time turned into action.

Scenario B, A contested public decision

A polity must decide a charged question, say, the siting of critical infrastructure. M1 and M3 ensure the claims in circulation are sourced and weighted; M2 renders the evidence map, separating the empirical disputes (which more data could settle) from the value differences (which it cannot). M5 maps the deliberation across thousands of participants, revealing that the genuine disagreement reduces to one or two cruxes rather than the all-encompassing conflict the unaided interface perceives. M6 lets decision-makers and the public see the same shared picture, and M7 helps each participant engage on their own terms. The decision that emerges is not necessarily unanimous, but it is made with a shared, legible understanding of what was actually at stake.

Scenario C, A single reader, a single article

Not every workflow is civilizational. A person reads an article. Their Companion (M7), querying the commons, surfaces through Rendering (M6) the article's provenance (M1) and the contextual trust weight (M3) of its sources; flags, via M1, an embedded image as machine-generated; and offers, from the Synthesis Engine (M2), the broader state of evidence on the article's central claim, including where it is contested. The reader is not told what to think. They are given an interface in which the trust judgment their biology can no longer make alone becomes, once again, something they can simply perceive.

These three scenarios span four orders of magnitude — a planet, a polity, a person — and that range is the point. The same modules, in the same arrangement, serve a global outbreak response and a single reader's afternoon. This is what distinguishes an interface from an application. An application solves a bounded problem for a defined user; an interface changes the terms on which its bearer meets any situation in its domain. The architecture is not eight tools that happen to share a backend. It is a single perceptual layer whose components recombine to fit whatever the bearer is facing, because the underlying need — to perceive verified structure, synthesized honestly, rendered where attention lives, governed against capture — is the same whether the stakes are civilizational or quietly personal.

It is also worth noting what does not change across the scenarios: the order of operations. In every case, nothing reaches perception that has not first acquired provenance, and nothing is rendered that governance does not constrain. The foundation band is not skipped when the matter is urgent, and the stewardship frame is not relaxed when the user is only one person. This invariance is deliberate. The moments when a system is most tempted to cut corners — the emergency, the trivial case — are exactly the moments when cutting them does the most damage, because they are the moments no one is watching closely. An interface that honoured its principles only when convenient would not be an interface its bearers could rely on; it would be one more thing to second-guess. The discipline of the architecture is that it holds its shape regardless of scale or stakes.

05 · Pathway

Adoption Pathway

A system of this scope is not deployed all at once, and it should not be. The architecture is designed so that early modules deliver standalone value before later ones exist, and so that adoption can follow the familiar shape of any genuine infrastructure: a slow, unglamorous foundation phase, a steep period of compounding returns once the pieces interlock, and a long maturation as the system becomes ambient.

Adoption S-curve across three phases A logistic adoption curve divided into foundation, compounding, and maturation phases. value / reach time → Foundation M1 · M3 provenance & trust Compounding M2 · M4 · M5 synthesis interlocks Maturation M6 · M7 ambient interface M8 governance present from the first day, not bolted on at maturity
Figure 11. An illustrative adoption pathway. The foundation modules (provenance, trust) deliver value on their own before the cognition layer exists; returns compound steeply as synthesis, detection, and deliberation interlock; and the system matures into an ambient interface. Governance is built in from the first day rather than retrofitted. Illustrative; the curve expresses intended sequencing, not a forecast.

The sequencing matters for a reason beyond convenience. Provenance and trust (M1, M3) are useful even to a single newsroom or research group with no other module present, which means the foundation can be built and validated by communities that benefit immediately, rather than requiring an act of faith in the whole system. Only once that verified substrate exists does synthesis (M2) have trustworthy material to work on, and only once synthesis exists does large-scale deliberation (M5) have a shared map to deliberate against. Each phase earns the next. Governance, uniquely, cannot wait its turn: a system that adds governance only at maturity has, by then, already shaped its incentives and accreted its power. M8 is present from the first commit.

There is a second reason the sequencing follows this shape, and it is about trust in the system itself rather than its technical dependencies. A perceptual prosthetic asks for an extraordinary thing: that people allow it to shape what they perceive. That permission cannot be demanded; it can only be earned, slowly, through demonstrated reliability in lower-stakes settings before the system is trusted with higher-stakes ones. Beginning with provenance and trust in bounded communities — a newsroom, a research group, a profession — lets the system accumulate a visible track record exactly as the trust graph (M3) would have it: contextually, on the basis of performance, open to inspection. By the time the interaction band reaches into everyday perception, the system has a history a reasonable person could examine. A version that launched fully formed and asked for that trust up front would, and should, be refused.

The compounding phase is where the architecture's central bet either pays off or fails. The claim is that the modules are worth dramatically more together than apart — that verified provenance makes synthesis trustworthy, trustworthy synthesis makes deliberation substantive, substantive deliberation makes rendering worth perceiving. If that interlocking value is real, adoption accelerates as each module amplifies the others. If it is not — if the pieces turn out to be merely additive rather than multiplicative — the steep middle of the curve never arrives, and the honest conclusion would be that the system is a collection of useful tools rather than a coherent interface. The architecture is structured so that this bet is testable early and cheaply, in the foundation phase, rather than discovered expensively at the end.

06 · Honesty

Risks, Failure Modes, and Limits

A vision paper that lists only benefits is itself a kind of fabrication, exactly the thing this architecture is meant to guard against. The most serious risks are worth naming plainly.

The capture risk

The gravest danger is the one Module 8 exists to address and can only mitigate, not eliminate: a system that shapes collective perception is the most valuable possible target for capture. Distributed stewardship raises the cost and visibility of capture, but no governance design is permanently capture-proof. This risk must be treated as ongoing rather than solved, and the system's own transparency is the primary defence, a captured commons that could be audited would, in principle, reveal its capture.

The new-orthodoxy risk

A system that synthesizes "what is known" could ossify into an instrument that defines a permitted consensus and marginalizes the dissent that, in science and in society, is sometimes right. The architecture's defences, preserving contested and unknown regions honestly (M2), modelling trust as contextual and decaying rather than fixed (M3), and surfacing rather than suppressing cruxes (M5), are deliberate countermeasures, but the temptation to collapse the map into a single answer will be constant and must be constantly resisted.

The dependence risk

A prosthetic can atrophy the faculty it supports. If the system does people's perceiving for them, it could erode the very capacities it was meant to extend. This is why the Companion (M7) is specified to scaffold understanding and raise capacity rather than to hand down conclusions, but the pull toward convenient dependence is real, and the design must keep choosing the harder path.

The hard limits

Finally, some limits are not risks to be managed but boundaries to be respected. The system cannot resolve genuine value disagreements; it can only clarify them, and a society that mistakes clarified values for a technical problem with a technical answer will misuse the tool. It cannot manufacture trust where the underlying conditions for trust are absent. And it remains, however sophisticated, an interface, a better map, never the territory. Hoffman's own framework insists on this: there is no final interface that shows reality as it is. There are only interfaces better or worse suited to the environment we must survive in. Commonsent aims to be a markedly better one. It does not pretend to be the last.

Capability comparison across perceptual dimensions Grouped bar chart comparing the evolved headset and Commonsent across five perceptual capability dimensions. low capability → high Visible threats Statistical risk Trust at scale Large-group coord. Long-horizon causality evolved headset alone with Commonsent
Figure 12. Where the second interface adds capability. The evolved headset (brown) excels at visible, proximate threats and little else in the modern environment; Commonsent (teal) is designed to lift the dimensions where biology is nearly blind. Note that on visible threats the system adds nothing, the prosthetic complements the biology rather than replacing it. Illustrative; bars are qualitative, not measured.
07 · Conclusion

Conclusion

The problems that most threaten the present, epistemic collapse, coordination failure, the erosion of shared reality, the silent approach of systemic risk, are often treated as separate crises with separate fixes. Seen through the interface lens, they are facets of one condition: a species-level mismatch between the perceptual apparatus evolution gave us and the environment we have built. The headset that served the savanna renders the networked world wrong, with perfect confidence.

The eight modules described here are not eight products. They are the parts of a single prosthetic, each closing one of the specific gaps between what the evolved interface can perceive and what the present demands. Provenance and trust rebuild the foundations of credibility that fabrication destroyed. Synthesis and detection extend perception into complexity and time that no individual mind can hold. Deliberation, rendering, and the companion connect that extended perception back to human action, individually and at scale. And governance keeps the whole apparatus from becoming the very thing it was built to defend against.

The interface we wear decides the world we can act in. We did not choose the first one. We can choose the second.

None of this is a claim that technology will save us, or that a better interface is sufficient for a better world. The architecture's own honesty section insists otherwise: it cannot resolve our disagreements, manufacture trust from nothing, or show us reality unmediated. What it can do is more modest and more urgent, give a civilization a perceptual interface fit for the environment it actually inhabits, rather than the one its biology was built for. That is what every consequential tool has done at the moments that mattered: not decorate existing capability, but extend the boundary of what can be perceived, and therefore thought, and therefore done.

The headset we wear shapes the world we can inhabit. The first one is no longer enough. It is time to build the next.

References & Further Reading

Hoffman, D.D. (2019). The Case Against Reality: Why Evolution Hid the Truth from Our Eyes. W.W. Norton., Hoffman, D.D., Singh, M., & Prakash, C. (2015). The Interface Theory of Perception. Psychonomic Bulletin & Review., Dunbar, R. (1992). Neocortex size as a constraint on group size in primates. Journal of Human Evolution., Ostrom, E. (1990). Governing the Commons: The Evolution of Institutions for Collective Action. Cambridge University Press., Henrich, J. (2015). The Secret of Our Success. Princeton University Press.