NicholicIndependent

Research note

Aver: The First AI Safety Research at Nicholic

Introducing Aver, the first AI safety research project at Nicholic: an assistant that shows where an answer came from, and what is disputed about it.

Published · Amended

Introducing Aver, the first AI safety research project at Nicholic: an assistant that shows where an answer came from, and how contested those sources actually are.

None of it has been run. The design is written and a first pass at the code is written, and this note describes both — but no question has ever been put to it and no answer has ever come out of it. That belongs at the top rather than in a closing paragraph, because a note that saves its status for the end has already misled everyone who stopped reading early.

What follows is the argument, the mechanism, and the parts that are known to be weak.

What it is meant to do

Most assistants hand back one confident answer and quietly discard the argument behind it. The reader never learns which claims were settled, which were disputed, or who disputed them. Disagreement is treated as noise to be cleaned up before it arrives.

Aver treats disagreement as part of the answer.

Underneath it sits Touchstone, a credibility engine built on two ledgers that stay permanently separate:

  • The accuracy ledger — falsifiable, checkable facts about a source's track record. Corrections issued. Retractions. How far a document sits from the primary source it describes. Of the claims that later resolved one way or the other, how many held.
  • The reception map — what public communities think of that source. Measured, displayed, and never acted upon.

The whole design rests on those two never being collapsed into one number, and the rest of this note is largely about why.

Four ways sourcing fails now

Laundered confidence. A model reads five sources of wildly different quality, synthesises them into fluent prose, and delivers the result in one even register. Nothing in the output distinguishes three peer-reviewed studies agree from one blog post asserted this.

Invisible curation. When an assistant declines to cite something, the reader is not told it happened. A source that was withheld and a source that was never found look identical from the outside.

Consensus flattening. Live disputes get resolved into a single narrative — usually whichever position was more common in the training data. Contested science, contested history and contested politics all receive the same smoothing.

Nothing to audit. There is no way to inspect why an assistant trusted what it trusted, no ledger to check, and no record when the behaviour changes.

The obvious fix, and why it fails

The intuitive repair is: let communities say which sources are unreliable, and stop citing those. It does not work, and the reason is worth being precise about, because it is not obvious from the outside.

Every major source is called unreliable by someone. One outlet by one crowd, another by a different crowd. Public health bodies, wire services, preprint servers, the large medical journals — all contested somewhere, by someone with a sincere argument. Applied consistently, a rule that drops contested sources does not produce a truthful assistant. It produces one that can only cite material too obscure for anyone to have argued about yet.

Popularity measures belief, not accuracy. Three different platforms will hold three different aggregate views of the same outlet, and none of the three is a measurement of whether that outlet got the facts right. A community's confidence is a fact about the community. It is worth knowing, and it is worth showing. It is not a verdict.

Any score that gates access becomes a target. The moment it is understood that organised downvotes remove a source from an assistant, organising downvotes becomes a strategy. What has been built at that point is a tool for quietly removing sources, with the keys handed out for free.

So a system designed to find truth through community signal, implemented in the obvious way, ends up better at suppressing sources than the thing it set out to replace — while wearing the badge of the opposite intention the entire time.

Surface the dispute, do not remove the source

The reframe the rest of the design is built on is a single line: surface the dispute, never suppress the source.

What a reader actually needs is not this outlet is unreliable, so it has been hidden. It is closer to this:

Same information. Nothing suppressed. More is known about the claim, and less has been decided on the reader's behalf, than the drop-the-contested rule delivers.

Why this is safety work

It would be easy to read the above as being about accuracy alone. It is not. Aver is a safety project, and the two failures it is built against are both safety failures — they simply fail in opposite directions.

The first is the obvious one. An assistant that is fluent and confidently wrong about a dose, a drug interaction, a filing deadline or whether a wall is load-bearing can injure someone, and the register it uses to be wrong is identical to the register it uses to be right. That is why the tier is assigned on the consequence of being wrong, and why the top tier refuses to blend conflicting sources into a single reassuring paragraph. Averaging two incompatible answers about a dose produces a third answer that nobody checked.

The second is less discussed and matters at least as much. A system that quietly withholds sources under pressure has taken the power to decide what is knowable and moved it somewhere nobody can see. It fails silently, it looks like working software the entire time, and the people affected by it are the least likely to be able to detect it. Building a credibility score that gates what a reader is allowed to see, and then defending it against manipulation, is a harder problem and a worse design than removing the reason to manipulate it. The cardinal rule is a safety mechanism before it is an epistemics one.

Three commitments follow from taking both seriously at once.

Safety measures are tiered, never global. A blanket restriction that slows and hedges every question in order to protect the small fraction that are genuinely high-stakes is a failure of the principle rather than a fulfilment of it. It also trains people to route around the assistant, which makes them less safe, not more.

The bar for withholding is demonstrable serious harm. Not discomfort, not unpopularity, not contestedness. This is controversial is never a reason to withhold something. This provides meaningful operational uplift toward mass casualties is. The distance between those two sentences is enormous, and Aver is designed to live inside it rather than to collapse it for convenience.

Customisation has a floor. A person can tune how much the assistant hedges, exclude sources, and choose which communities' reception is shown. A person cannot switch off provenance, cannot disable the top tier's handling, and cannot drop below the evidence floor at the two highest tiers — because those protect people who are not in the room to argue for themselves.

One caveat belongs here rather than in a footnote, and it is the same one the rest of this note carries. None of these are running controls. They are design commitments, written down before the code that would implement them has been run even once, and the published commitments on this site set a deliberately high bar for calling something a safety property: it does not count until it exists and is switched on. Nothing in this section clears that bar yet. It is stated as intent, in public, dated, so that it can be checked against whatever actually gets built — which is the only reason writing it down this early is worth anything.

Touchstone: the two ledgers

The accuracy ledger

Admissible inputs, all of them falsifiable and all of them evidence-linked:

  • Published corrections, from the outlet's own corrections page.
  • Retractions — the strongest available signal.
  • Primary-source distance: did the piece link the study, or rewrite a press release?
  • Resolved-claim hit rate: of claims that later resolved definitively, what fraction held.
  • Methodology disclosure, which is binary and checkable.
  • Correction latency and prominence, measurable in days and in page position.

Inadmissible, permanently:

  • Bias ratings from third-party organisations.
  • Expert panel trust scores.
  • Public trust surveys.
  • Political lean classifications.
  • Any reception signal whatsoever.

Those are opinion wearing a lab coat. Importing one is how a truth engine quietly turns into an ideology engine, and it happens gradually enough that nobody notices the transition while it is happening.

The ledger's output contract is a triple, never a bare number:

(score: float | None, confidence: float, evidence: list[Evidence])

A score of None at low confidence is a common and correct result. Most sources have insufficient evidence for a long time, and saying so plainly beats inventing a number the evidence cannot support. Every event carries an inspectable address, enforced by a database constraint rather than by convention: an accuracy claim that cannot be shown to someone is not an accuracy claim, it is an opinion with a number attached to it.

The reception map

What it measures is how a source is received across public communities, weighted by how expensive the account making each observation would be to fake.

What it reports:

  • Sentiment by community, never a global average. Broadly positive in one community, broadly negative in another is informative. The mean of those two is noise.
  • Dispersion — is criticism spread across many communities and long stretches of time, or concentrated in one place over six hours?
  • Burst signature — five hundred comments from four hundred accounts created this month has a completely different shape from five hundred spread across two years.
  • Weighted volume — account age and how many distinct communities someone posts in feed a per-observation weight. A three-day-old account posting in one place counts for almost nothing; a four-year-old account active across twenty communities counts for one.

The weighting method is published rather than hidden. An attacker knowing it does not make evasion cheap: evading it requires aged, diverse accounts, which is exactly the expense the weighting exists to force.

There is a deliberate absence in the data model here. No single overall sentiment number exists anywhere — not as a column, not as a property, not as a computed field. Providing one would invite precisely the merge described below, and something that can be sorted by eventually gets sorted by.

The cardinal rule

This is an architectural constraint rather than a policy setting, and in the written code it is a test that parses the source tree and fails the build if the retrieval module ever imports the reception module, or if a reception value ever turns up in a condition, a loop test, a comprehension filter or a sort key. The test reads the syntax tree rather than the raw text, so the prose explaining the rule is free to discuss it openly without tripping over itself.

The reason this matters more than it appears to: it removes the incentive for manipulation rather than trying to defeat the manipulation. If ten thousand coordinated downvotes cannot remove a source, nobody profits from organising ten thousand coordinated downvotes. Nothing has to be spent defending against brigades, because there is no prize at the end of one.

Every other mitigation in the design is secondary to that one.

Why the separation is binary

Merging the two ledgers is tempting. One number is easier to display, easier to sort, and easier to explain to someone in a sentence.

The moment a reception signal carries any weight in a score that affects what the reader sees, popularity begins standing in for truth — silently, with no error message, in a way that is close to impossible to detect from outside the system. There is no ten-percent reception weighting that is safe. The separation is binary, and the code is written so the two ledgers meet only as separate fields on the same object, never as inputs to a shared calculation.

Writing combined = 0.7 * accuracy + 0.3 * reception, at any weighting, rebuilds the thing this project exists to replace.

How one question gets answered

In order:

  1. The question arrives at the console.
  2. A stakes tier is assigned — cheap, local, and before retrieval, so it can size what follows.
  3. Search and fetch. No reception signal reaches this step or anything downstream of it.
  4. Primary-source elevation: documents that are the original record rank above coverage of that record.
  5. Profile assembly. Both ledgers are attached to each document, kept apart.
  6. Synthesis. The model receives the documents, both profiles and the tier, under a strict output contract with separate sections for the answer, for what could not be established, and for where the documents disagree.
  7. A verification pass at the top two tiers, checking that each factual sentence maps to a cited document.
  8. Render: the answer with inline references, then a source panel showing both ledgers side by side.
  9. Log everything retrieved — including anything the reader's own filters removed.

Reception reaches the model labelled explicitly as context for the reader and not as an instruction. That framing is load-bearing. Models are agreeable, and if it slips even slightly they begin helpfully avoiding sources with poor reception, which reintroduces community-driven removal through a side door where no test is watching.

Stakes, not sensitivity

Tiers are assigned on the consequence of being wrong:

  • T0 — trivia, entertainment, creative, opinion. Answer freely, cite where useful.
  • T1 — general factual, history, science explanation, technical. Cite substantive claims, flag conflicts.
  • T2 — health, legal, financial, safety-relevant. Require a primary source or say plainly that none was found, and surface every conflict prominently.
  • T3 — dosing, drug interactions, legal deadlines, structural and electrical safety. Everything T2 does, plus a refusal to synthesise across conflicting sources: each position is presented separately, the conflict is named, and qualified human review is recommended.

The distinction that carries the weight: politically uncomfortable is T0 or T1, not T2. Discomfort is not danger. Conflating those two is the most common way a system built to prevent harm drifts into removing things it was never meant to touch, and tiering on the consequence of being wrong — never on who might object to the answer — makes that conflation structurally hard rather than merely discouraged.

That is the same idea TruXSocial runs on in a different product: name the harms in advance, in public, keep the list short, and refuse to stretch it.

Primary-source distance

Every retrieved document gets a distance from the primary source of its central claim:

  • 0 — the document is the primary source: the study, the filing, the transcript, the dataset.
  • 1 — links and quotes the primary source directly.
  • 2 — cites the primary source without linking it.
  • 3 — reports on someone else's reporting.
  • 4 — no attribution chain that resolves.

Distance-0 documents are elevated above the coverage of them, always. An outlet reports that the agency found X becomes the agency found X, and here is that outlet's framing of it.

This is the cheapest high-value move in the whole system. For a large share of factual questions it steps around the argument about outlet reliability entirely, because the argument was never really about the outlet. Nobody brigades a court transcript.

What the score refuses to say

The accuracy score is a weighted sum of four components, with a penalty:

accuracy = w1*transparency + w2*correction_quality
         + w3*hit_rate + w4*distance - retraction_penalty

  w1 = 0.20   transparency        does the outlet publish corrections at all
  w2 = 0.20   correction quality  how quickly, and how prominently
  w3 = 0.35   hit rate            of claims that resolved, what fraction held
  w4 = 0.25   distance            does it link its sources or paraphrase blind

confidence = 1 - exp(-n_eff / 12)   computed from evidence VOLUME, and
                                    deliberately independent of the score

Every number there is a starting guess rather than a derived value, and they are grouped in one place in the code so they can be fitted against a labelled evaluation set later, instead of pretending to a precision they have not earned.

Four properties matter more than the numbers themselves.

Confidence is not part of the score. A high score built on two data points is not a high-confidence high score, and keeping the two quantities apart is what allows the system to say so.

Below the display threshold, no number appears at all. The reader sees insufficient record, with the effective evidence count beside it. Showing a number the evidence cannot support is a lie with extra steps, and an honest gap renders as information rather than as an error.

Correction quality is a positive term, on purpose. If corrections only ever lowered a score, the system would punish outlets that make checkable claims and reward outlets vague enough never to be caught being wrong. That perverse incentive is very easy to build by accident, and prompt, prominent corrections should raise a source's standing relative to silence.

A missing component is renormalised away rather than read as zero. A source with no correction history is not scored as though it corrects nothing. It is scored on what is actually known about it, across the weights that apply.

Evidence decays: a half-life of three years for accuracy, one year for reception. A scandal from a decade ago should not permanently condemn a newsroom that rebuilt itself, and community sentiment is more volatile and less informative over time than a correction record is.

Scoring happens at the level of a source and a beat, never at the level of an outlet. Asking whether a large newsroom is reliable is close to meaningless: its election desk, its health desk and its opinion pages have entirely different records, and the schema is shaped so that outlet-global scoring is awkward to write rather than merely discouraged.

The reader's own filters, said out loud

A person can exclude sources, adjust how much the assistant hedges, change flagging thresholds, and choose which communities' reception is displayed. What a person cannot do is switch off provenance or drop below the evidence floor at the top two tiers, because those protect people who are not in the room.

Every answer after an exclusion says what the exclusion removed. Without that line a person builds a filter bubble in the first week and forgets they built it. The query log records the full retrieval set including everything the filters dropped, so the record outlives the session that created it.

A silent personal filter is the same failure as a silent global one. It is just harder to notice, which arguably makes it worse.

What it runs on, and what it keeps

Aver runs locally, on one consumer graphics card, with inference handled by a local server so models can be swapped with a single command. The reference target is an 8B model with a generous context window rather than a larger model squeezed into a small one — every query carries five to eight full documents into the prompt, so room to read matters more than parameter count for the job actually being done. Rough budget on a 12 GB card:

model weights, 8B at Q5_K_M          ~5.8 GB
KV cache, 16k context, Q8            ~1.3 GB
embedding model, on CPU               0.0 GB
headroom and desktop compositor      ~1.5 GB
────────────────────────────────────────────
total                                ~8.6 GB

The model is a commodity in this architecture. The credibility layer is the product, and the reason for holding the output format stable across model swaps is that the reasoning engine stays replaceable.

What is kept, and where: the query log lives in a single database file on the person's own machine. Nothing is sent anywhere to be analysed. The reception map, when it is built, is designed to store permalinks and derived metrics rather than copies of what people wrote, and its weighting inputs — how old an account is, how many communities it posts in — are derived and non-identifying, which is a property to preserve rather than a coincidence to rely on.

What does leave the machine is the search queries and the page fetches, because retrieval is retrieval. Saying so is cheaper than being caught not saying it.

What is not built

Set against that design, the honest state of the work:

  • Reception ingestion is not built. Nothing is collected from anywhere. The module returns an empty record and the interface reads no reception data.
  • The corrections crawler is not built. The accuracy ledger has a schema and a scoring function and no events in it, so every source reports insufficient record — which is, at the moment, the true answer.
  • Claim-level scoring is not built, and it is research rather than a task with a known shape.
  • The stakes classifier is a list of text patterns, standing in for the trained classifier meant to replace it.

And the line that matters most: none of this has been run. The database has never been created, the console has never been opened, no question has been asked of it and no answer has come back. The tests described above have never been executed either. It is a design and a first draft of an implementation, published in that state deliberately.

The limits worth naming before somebody else finds them

The ledgers stay thin for a long time. For months most sources will return insufficient record. That is correct behaviour, and it also makes an early version feel empty. Seeding from public, curated source-reliability corpora and from public retraction databases is the mitigation, along with rendering the gap as information rather than as a failure.

The pattern-matching tier classifier will misfire in both directions. A question about a poison scene in a film trips the health tier; a genuine question about a safe amount of a common painkiller can fall through to the general tier if it avoids the obvious words. Both directions are wrong, and only one of them is dangerous.

A small local model under a strict output contract will sometimes break format. The intended handling is to validate, retry once, and then fall back to showing the retrieved documents rather than a confidently wrong answer.

Corrections pages have no standard. Every outlet is a small bespoke parser, which is unglamorous, slow, and the highest-value contribution anyone could make to the project.

Beat classification is genuinely fuzzy. A story about vaccine policy is health and politics at once. Multi-beat tagging with partial weights is the likely answer, and it complicates every scoring query, so it is marked unsettled rather than quietly assumed.

Accusations of bias arrive anyway. Touchstone will sometimes report a poor accuracy record for a source that some community holds in high regard, and publishing the full evidence-linked ledger is the only real defence. It is a good one, but only if the ledger is complete, timestamped and public before the accusation arrives — which is an argument for building the transparency surface early rather than at the moment it is needed.

Two questions the design has not settled. Whether fact-checker rulings are admissible: they are documented and evidence-linked, which argues for, but they involve editorial judgement about which claims get checked at all, which is a reception-like signal in disguise. The current position is to exclude them from scoring and permit them as a displayed source like any other. And which licence the code carries, where the case for a strongly reciprocal one is that a closed fork quietly reintroducing reception-based filtering would be the worst available outcome.

The order the work happens in

Skeleton first: schema, local inference, a console that answers with citations, and the cardinal-rule test passing. Then the accuracy ledger — corrections crawling for a set of seed outlets, primary-source distance from link structure, confidence gating, a trained tier classifier, and an evaluation set of hand-labelled questions with known-good answers. Then the reception map. Then a public transparency surface where an outside critic can audit any score shown. Only then claim-level scoring.

The order is not arbitrary, and reordering it is the mistake to avoid. The first two stages depend on nobody's permission. Reception depends on third parties who can revoke access at any time, and on platforms whose research access has already been closed or narrowed — one large platform shut its public research tool in 2024 with no replacement, so it is out of scope rather than promised. Building the half that is under one person's control first is also building the half that is about truth rather than belief.

No dates are attached to any of it, for the reason the refused promises already give: a date given now would be a guess presented as a plan.

Why publish a design for something that has not run

Because the parts of this worth holding to are easiest to write down before there is anything to lose by writing them down.

A rule that reception never gates retrieval costs nothing today. It costs something on the day a source with a genuinely bad record is also loudly disliked, and the easy thing is to let the second fact quietly do the work of the first. Publishing the rule while it is still free is what makes it checkable later — and this note can be held against whatever gets built, including by the person who wrote it.

The same argument produced the commitments on this site before there was a product to break them on. This is that method applied to a piece of research, and the status section above is the first thing to check when there is finally something running to compare it against.

Share this note

Where the browser offers a sharing sheet, this opens it, with a story-sized image of the note attached where the browser allows. Otherwise it opens that image, which can be saved.