Every number on ProConWiki comes from a transparent formula. This page explains the evidence-strength score on each argument card, the Supports/Against balance on each claim, and the composite review score — what each measures, how much each factor weighs, and what we deliberately leave out. Updated 2026-07-31: the claim balance now scores each side the way it scores each argument — strongest counts in full, the rest add 10% each — so argument count can no longer move a side. The per-argument formula is also stated exactly as implemented (secondary sources contribute quality-scaled).
Each argument shows a score like 72/100. That number measures how good the sources behind the argument are — not how popular it is, and not whether we think it wins.
A peer-reviewed study counts a lot. Government data counts a lot. A company report counts less. A news story counts less than that. Newer sources count a little more than old ones. The best source matters most — adding lots of weak sources barely moves the score.
The label next to the score (like "Direct evidence" or "Logical inference") tells you how the argument uses its sources: did someone measure this directly, or is it a reasonable conclusion drawn from the facts?
At the top of a claim you see something like Supports 64% · Against 36%. Each side is measured the same way an individual argument is: the strongest argument on that side counts in full, and every other argument on that side adds 10%. It is weighted by evidence quality, not votes.
Why not simply add the sides up? Because then splitting one argument into two would inflate a side without adding a single new source. Under this rule the count barely moves the number and the quality decides it — the same "best one dominates" principle used inside each argument card.
Two honest notes: arguments that say "it’s complicated" (nuanced) are shown separately and are not inside this split. And the balance describes the evidence we found — it is not a promise about every study that exists, and it is not a probability that the claim is true.
Used when an argument is re-scored in the reasoning-graph review process. The number shown on argument cards is the evidence strength described above.
Is the evidence any good? A scientific study counts more than a blog post. A study that other scientists repeated and got the same result counts even more. This is the biggest part of the score because good evidence is the foundation of a good argument.
How many separate sources back this up? One study is a start. Ten studies saying the same thing is much stronger. But quantity alone isn't enough — ten bad studies don't beat one great one.
Does the argument actually make sense? Even with great evidence, the reasoning has to connect. If someone says "the sky is blue, therefore pizza is healthy," the evidence is fine but the logic is broken.
Has this argument survived being challenged? When someone argues against it and the argument still holds up, that makes it stronger. An argument nobody has tried to knock down isn't as proven as one that's been tested and survived.
What's the track record of the person who contributed this argument? If they've contributed good arguments before that held up over time, their new arguments get a small boost. But only a small one — evidence matters far more than reputation.
Are the sources actually different from each other? If five articles all got their information from the same original study, that's really just one source. Independent sources that reached the same conclusion separately are worth more.
How many people liked or upvoted an argument does not change its score. Ever. You can see how popular an argument is, but that number has zero effect on ranking. This is on purpose. If popularity counted, the most-liked argument would always win — even if the evidence said otherwise. ProConWiki protects arguments that have strong evidence, even when most people disagree.
Why do we show the formula? Because a platform that claims to be about transparent reasoning should be transparent about how its own reasoning works. You deserve to know exactly why an argument scored the way it did.
The score on each argument card (e.g. 72/100) is its evidence strength: a measure of the quality of the sources it cites, computed the same way for every argument:
1. Source type sets the base. A meta-analysis starts near the top (0.95), a randomized trial 0.90, a peer-reviewed paper 0.80, government data 0.75, an institutional report 0.70, expert opinion 0.55, news reporting 0.45, an industry report 0.40, opinion pieces 0.25. A source we cannot classify scores 0.05 — being unclassifiable is treated as a defect, not a pass.
2. Modifiers adjust it. Recency (newer counts more), independence of the source, replication status, and sample size each nudge the base up or down.
3. Relevance weighs it. Each source is weighted by how directly it bears on this specific argument.
4. The best source dominates. The final score is the strongest source’s value plus 10% of each additional source (capped at 100). Ten weak citations cannot beat one excellent study, and padding a card with extra links barely moves the number.
The reasoning label next to the score (Direct evidence, Data analysis, Expert opinion, Logical inference) describes how the argument connects its sources to its conclusion. It is displayed for your judgment and does not change the number.
Source types are classified by our research system, not hand-entered. If a classification looks wrong, use "Help improve this argument" — misclassification is a bug we want to hear about.
The Supports / Against split takes each side's strongest argument at full weight, adds 10% of every other argument on that side, and normalizes the two. It is weighted by source quality — never by votes or popularity, and never by how many arguments a side happens to have.
Three things this number is not: it does not include nuanced arguments (they are a separate category shown on the page — often a large share of the total weight); it reflects the evidence gathered for this page, not every study in existence; and it is not a probability that the claim is true. It answers one question: of the directional evidence on this page, how does the quality-weighted mass divide?
Used when an argument is re-scored in the reasoning-graph review process. The number shown on argument cards is the evidence strength described above.
Assesses the credibility and strength of the evidence cited. Primary sources score higher than secondary reporting. Peer-reviewed research scores higher than opinion pieces. Studies that have been independently replicated receive the highest quality ratings. This is the largest factor because the quality of evidence is the most important determinant of argument strength.
Counts the number of independent sources supporting the argument. Multiple sources converging on the same conclusion provides stronger support than a single source, even a high-quality one. However, quantity is weighted less than quality — ten weak sources do not outweigh one rigorous study.
Evaluates whether the reasoning is valid. Checks for common logical fallacies (ad hominem, false dichotomy, appeal to authority, etc.), whether conclusions follow from premises, and whether the argument addresses the actual claim rather than a strawman. An argument can have excellent evidence but poor logic if the evidence doesn't actually support the conclusion being drawn.
Measures how well an argument holds up when challenged by opposing arguments. An argument that has faced strong counterarguments and survived scores higher than one that hasn't been tested. This rewards resilience — the intellectual equivalent of stress-testing. Arguments that have been refuted or significantly weakened by counterarguments see their counterpressure score decline.
Reflects the track record of the person who contributed the argument. Contributors whose past arguments have held up well over time (maintained or improved their scores) receive a modest credit. This creates an incentive for thoughtful, well-evidenced contributions. The weight is deliberately low — what matters is the argument itself, not who made it.
Evaluates whether the cited sources are genuinely independent of each other. Five news articles all citing the same original study represent one independent source, not five. Sources that arrived at similar conclusions through separate research, data collection, or analysis receive higher independence scores. This prevents citation cascades from artificially inflating evidence quantity.
Upvotes, likes, and popularity signals are displayed on every argument but are excluded from score computation. This is a hard architectural rule, not a tunable parameter. The moment popularity influences scoring, the platform becomes a consensus engine rather than a reasoning engine. ProConWiki is designed to protect the well-evidenced minority position — the argument that has strong evidence even when most people disagree. History repeatedly shows that majority opinion and evidentiary strength are different things.
A platform that claims transparent reasoning should be transparent about its own reasoning methodology. The weights above are public by design. The specific algorithms that compute each factor (how contributor credit is calculated, how counterpressure resilience is measured, how source independence is detected) are proprietary — but the framework and its priorities are fully visible.
Displayed per argument as round(evidence_strength_score × 100). For each cited evidence row:
quality = base(source_type) × recency(year) × independence × replication × sample
Base table: meta_analysis 0.95 · randomized_controlled_trial 0.90 · peer_reviewed 0.80 · government_data 0.75 · institutional_report 0.70 · data_analysis 0.65 · expert_opinion 0.55 · news_report 0.45 · industry_report 0.40 · opinion_editorial 0.25 · anecdote 0.15 · unverified 0.05. Modifiers are bounded multiplicative adjustments (recency 0.6–1.0; independence, replication and sample-size each defined per source class).
Per-argument aggregation over its evidence set, with each item weighted by argument_relevance_score: sort quality × relevance descending, take the top five, then strength = top[0] + 0.1 × Σ(s² for s in top[1:5]), clamped to [0,1]. Secondary sources contribute quality-scaled, so a weak corroborating citation adds far less than a strong one. Design intent: best-source-dominates — citation padding is structurally ineffective.
Worked example: two peer-reviewed sources whose quality × relevance land at 0.68 and 0.63 → 0.68 + 0.1 × 0.63² ≈ 0.72 → 72/100. A third source at 0.30 would add only 0.009.
reasoning_type (direct_evidence | data_analysis | expert_opinion | logical_inference) is machine-assigned display metadata describing the inferential mode; it does not currently modulate the score. Source types are machine-classified from the citation (title/venue/URL); unclassifiable sources are deliberately scored 0.05 so vocabulary violations surface rather than pass silently.
Sort a side’s active arguments by evidence_strength_score descending, then side_weight = top[0] + 0.1 × Σ(rest); Supports% = pro / (pro + con). Nuanced arguments carry weight in the three-way distribution shown on the claim but are excluded from the two-way Supports/Against split.
This deliberately mirrors the per-argument rule one level up: best one dominates, the rest add a diminishing 10% each. It replaced a plain Σ over the side on 2026-07-31. Under a sum, argument count was a lever — splitting one argument in two raised a side’s mass without adding evidence. Under this rule that split yields 0.99 instead of 1.8, so quality decides the balance and structure does not. Unlike the per-argument score, the side weight is not clamped to 1.0 and has no top-five cutoff: it is only ever read as the ratio pro / (pro + con), where a ceiling would erase real differences rather than bound a defined range.
Epistemic scope: the balance is a statement about the quality-weighted evidence retrieved and curated for this claim, not an estimate of the full literature and not a posterior probability. Research deliberately searches both sides, but retrieval asymmetry is a real limit and we state it rather than hide it.
Applies to reasoning-graph re-scoring; the card number is the evidence-strength computation above.
Evaluates source credibility on a multi-tier hierarchy: peer-reviewed and replicated research at the top, followed by peer-reviewed but unreplicated, primary-source government or institutional data, expert analysis, quality journalism, and opinion or anecdotal sources at the bottom. Each piece of evidence cited by an argument is classified and the aggregate quality score is normalized to [0, 1]. The classification methodology is proprietary, but the hierarchy is designed to reward empirical rigor and primary-source proximity.
Counts independent evidence sources cited by the argument, with diminishing returns. The first source contributes the most; additional sources provide progressively smaller marginal gains. This reflects the epistemic reality that the jump from zero sources to one is far more significant than the jump from nine to ten. The precise diminishing-return curve is proprietary.
Assesses inferential validity: whether stated conclusions follow from cited premises, whether the argument addresses the claim directly (vs. strawman or tangential reasoning), and whether formal or informal logical fallacies are present. Detected fallacies reduce L(a) proportionally to severity. The fallacy detection and severity weighting methodology is proprietary.
Measures an argument's performance under adversarial scrutiny from opposing arguments. When strong counterarguments exist (high S for opposing arguments), an argument that maintains its evidence quality and logical validity scores higher on C. Arguments that have been effectively rebutted — where the counterargument's evidence directly undermines the original premises — see C decline. Untested arguments (no counterarguments exist yet) receive a neutral baseline rather than a high score. The resilience computation is proprietary.
Reflects the contributing user's historical argument quality. Computed from the contributor's past arguments' score trajectories: arguments that maintained or improved scores over time (indicating they held up as new evidence and counterarguments emerged) increase R. Arguments that degraded significantly decrease it. The weight is deliberately constrained to 10% to prevent reputation from dominating substance. The specific credit computation is proprietary.
Detects citation dependency chains and adjusts for them. If multiple cited sources trace back to the same original research, data set, or press release, they are clustered as a single effective source for quantity purposes, and the independence score is penalized. Sources that demonstrably conducted independent research, used separate data, or arrived at conclusions through different methodological approaches receive higher independence scores. The dependency detection methodology is proprietary.
Resonance signals (upvotes, likes, agreement counts) are captured and displayed as social metadata but are architecturally excluded from S(a) computation. This is enforced as a hard constraint, not a tunable weight. The exclusion is motivated by the well-documented divergence between popular agreement and evidentiary strength across domains including public health, economic policy, and scientific consensus formation. Inclusion of popularity signals would create a positive feedback loop where highly-visible arguments accumulate votes that increase their score, which increases their visibility, irrespective of their evidentiary merit. ProConWiki is a reasoning engine, not a consensus engine.
The weights and factor definitions on this page are public. The specific algorithms implementing each factor — how Q classifies source tiers, how C computes resilience under counterargument pressure, how R tracks contributor trajectories, how I detects citation dependency chains — are proprietary trade secrets. This separation is intentional: the framework and priorities are transparent, while the competitive implementation remains protected. A patent application covering the overall methodology has been filed.
The balanced synthesis displayed on each claim page is generated by AI from the scored argument set. The synthesis weighs arguments proportionally to their composite scores when summarizing the state of evidence. Arguments below a minimum score threshold are excluded from synthesis to prevent low-quality contributions from polluting the summary. The synthesis is regenerated when the argument set or scores change materially.
Merripedia ranks providers by behavioral trust — what they do, not what people say about them.
We watch how providers actually behave on the platform. Do they show up when they say they will? Do people come back to them? Do engagements complete successfully? These are the kinds of signals we use.
We deliberately exclude popularity metrics. There are no star ratings, no review scores, no like counts, and no way for social proof to influence ranking. This isn't a policy choice — it's an architectural one. Our system literally cannot factor in popularity because the data structure has no place for it.
The providers we recommend earned their ranking through reliable behavior, not through accumulating five-star reviews. A provider who has served three people exceptionally well can rank higher than one who has served a hundred people adequately.
Your ranking reflects your actual service quality. You don't need to ask clients for reviews, and no competitor can outrank you by gaming a rating system. Your trust score improves when you deliver reliably. It declines when you don't. It recovers when you return to form.
We do not publish the specific signals we measure, how they're weighted, or the thresholds that determine ranking. This is deliberate.
If we published the formula, bad actors would optimize for the metrics rather than for genuine service quality. A provider could game specific signals to build a score, then exploit the trust they manufactured. Keeping the formula private protects the integrity of the system — which protects you.
This is the same model used by credit scoring systems, search engines, and academic review processes: publish the principles, protect the implementation.
Principled transparency — explaining what the system values and why — serves users better than performative transparency that publishes formulas competitors and bad actors can exploit.
Some "Relevant Resources" links on claim pages are affiliate links (marked with an Affiliate chip): if you buy a book through one, ProConWiki earns a small commission from Bookshop.org at no cost to you. Every resource is hand-approved before it appears. Affiliate relationships never influence argument ranking, evidence scoring, or synthesis — the resource system is structurally separate from the scoring pipeline and cannot feed it.