Every number on ProConWiki comes from a transparent formula. This page explains the evidence-strength score on each argument card, the Supports/Against balance on each claim, and the composite review score — what each measures, how much each factor weighs, and what we deliberately leave out. Updated 2026-07-31: the claim balance now scores each side the way it scores each argument — strongest counts in full, the rest add 10% each — so argument count can no longer move a side. The per-argument formula is also stated exactly as implemented (secondary sources contribute quality-scaled).
Each argument shows a score like 72/100. That number measures how good the sources behind the argument are — not how popular it is, and not whether we think it wins.
A peer-reviewed study counts a lot. Government data counts a lot. A company report counts less. A news story counts less than that. Newer sources count a little more than old ones. The best source matters most — adding lots of weak sources barely moves the score.
The label next to the score (like "Direct evidence" or "Logical inference") tells you how the argument uses its sources: did someone measure this directly, or is it a reasonable conclusion drawn from the facts?
At the top of a claim you see something like Supports 64% · Against 36%. Each side is measured over its distinct sources: the strongest source on that side counts in full, and every other distinct source adds 10%. It is weighted by evidence quality, not votes — and because duplicate citations of the same source are counted once, repeating or splitting arguments cannot move a side at all. (Updated 2026-08-22: the unit changed from arguments to distinct sources, closing a measured loophole where duplicated arguments could still inflate a side through the 10% tail.)
Why not simply add the sides up? Because then splitting one argument into two would inflate a side without adding a single new source. Under this rule the count barely moves the number and the quality decides it — the same "best one dominates" principle used inside each argument card.
Two honest notes: arguments that say "it’s complicated" (nuanced) are shown separately and are not inside this split. And the balance describes the evidence we found — it is not a promise about every study that exists, and it is not a probability that the claim is true.
Used when an argument is re-scored in the reasoning-graph review process. The number shown on argument cards is the evidence strength described above.
Is the evidence any good? A scientific study counts more than a blog post. A study that other scientists repeated and got the same result counts even more. This is the biggest part of the score because good evidence is the foundation of a good argument.
How many separate sources back this up? One study is a start. Ten studies saying the same thing is much stronger. But quantity alone isn't enough — ten bad studies don't beat one great one.
Does the argument actually make sense? Even with great evidence, the reasoning has to connect. If someone says "the sky is blue, therefore pizza is healthy," the evidence is fine but the logic is broken.
Has this argument survived being challenged? When someone argues against it and the argument still holds up, that makes it stronger. An argument nobody has tried to knock down isn't as proven as one that's been tested and survived.
What's the track record of the person who contributed this argument? If they've contributed good arguments before that held up over time, their new arguments get a small boost. But only a small one — evidence matters far more than reputation.
Are the sources actually different from each other? If five articles all got their information from the same original study, that's really just one source. Independent sources that reached the same conclusion separately are worth more.
How many people liked or upvoted an argument does not change its score. Ever. You can see how popular an argument is, but that number has zero effect on ranking. This is on purpose. If popularity counted, the most-liked argument would always win — even if the evidence said otherwise. ProConWiki protects arguments that have strong evidence, even when most people disagree.
Why do we show the formula? Because a platform that claims to be about transparent reasoning should be transparent about how its own reasoning works. You deserve to know exactly why an argument scored the way it did.
The score on each argument card (e.g. 72/100) is its evidence strength: a measure of the quality of the sources it cites, computed the same way for every argument:
1. Source type sets the base. A meta-analysis starts near the top (0.95), a randomized trial 0.90, a peer-reviewed paper 0.80, government data 0.75, an institutional report 0.70, expert opinion 0.55, news reporting 0.45, an industry report 0.40, opinion pieces 0.25. A source we cannot classify scores 0.05 — being unclassifiable is treated as a defect, not a pass.
2. Modifiers adjust it. Recency (newer counts more), independence of the source, replication status, and sample size each nudge the base up or down.
3. Relevance weighs it. Each source is weighted by how directly it bears on this specific argument.
4. The best source dominates. The final score is the strongest source’s value plus 10% of each additional source (capped at 100). Ten weak citations cannot beat one excellent study, and padding a card with extra links barely moves the number.
The reasoning label next to the score (Direct evidence, Data analysis, Expert opinion, Logical inference) describes how the argument connects its sources to its conclusion. It is displayed for your judgment and does not change the number.
Source types are classified by our research system, not hand-entered. If a classification looks wrong, use "Help improve this argument" — misclassification is a bug we want to hear about.
Nobody at ProConWiki chooses what gets analyzed — not an editor, not the founder. The pipeline is mechanical and every number in it is published here:
1. Anyone suggests a topic. An AI gate checks one thing only: can this become a single, debatable claim that real-world evidence can support or oppose? Junk is declined politely; vague ideas come back with a sharper suggested framing.
2. The community votes. Every topic that passes the gate goes on the public Topic Queue — including reworded ones, shown under their sharper framing. One vote per person per topic. Votes lose half their weight every 14 days, so lasting interest beats a one-day spike. The queue holds the top 50 topics; a topic below the bar for 30 days retires (and can be re-suggested).
3. A published rule releases research. A topic is researched when its decayed vote score reaches 5.0 and it has been in the queue at least 3 days, within a budget of 2 promotions per day. Topics our gate estimates as narrow-audience need a higher bar (3× the score, 7 days) — friction against coordinated voting, applied by formula, not judgment. We also estimate each topic's audience breadth internally; we disclose that the estimate exists and don't publish per-topic values, so the estimate can't anchor the voting it exists to check.
4. The result credits its origin. A claim born from a suggestion shows who suggested it and how many votes promoted it.
Two honest notes. First: this design favors broad topics — a regionally- or locally-scoped question faces a higher effective bar today, and we say so rather than hide it. Second: news never rewrites an analysis directly. A news event can surface a claim and accelerate its evidence refresh, but the synthesis changes only when the scored evidence materially changes — headlines are weak evidence in the hierarchy above, and they stay that way no matter how loud the news cycle is.
These constants (5.0, 14 days, 3 days, top 50, 2/day…) are operating policy: they may be tuned as the site grows, and this page is always the current statement of them.
The Supports / Against split is measured over each side’s distinct sources: the strongest source on that side counts in full, every other distinct source adds 10%, and the two sides are normalized. It is weighted by source quality — never by votes or popularity, and never by how many arguments a side happens to have. (Updated 2026-08-22: the unit changed from arguments to distinct sources; see the version history below.)
Three things this number is not: it does not include nuanced arguments (they are a separate category shown on the page — often a large share of the total weight); it reflects the evidence gathered for this page, not every study in existence; and it is not a probability that the claim is true. It answers one question: of the directional evidence on this page, how does the quality-weighted mass divide?
Used when an argument is re-scored in the reasoning-graph review process. The number shown on argument cards is the evidence strength described above.
Assesses the credibility and strength of the evidence cited. Primary sources score higher than secondary reporting. Peer-reviewed research scores higher than opinion pieces. Studies that have been independently replicated receive the highest quality ratings. This is the largest factor because the quality of evidence is the most important determinant of argument strength.
Counts the number of independent sources supporting the argument. Multiple sources converging on the same conclusion provides stronger support than a single source, even a high-quality one. However, quantity is weighted less than quality — ten weak sources do not outweigh one rigorous study.
Evaluates whether the reasoning is valid. Checks for common logical fallacies (ad hominem, false dichotomy, appeal to authority, etc.), whether conclusions follow from premises, and whether the argument addresses the actual claim rather than a strawman. An argument can have excellent evidence but poor logic if the evidence doesn't actually support the conclusion being drawn.
Measures how well an argument holds up when challenged by opposing arguments. An argument that has faced strong counterarguments and survived scores higher than one that hasn't been tested. This rewards resilience — the intellectual equivalent of stress-testing. Arguments that have been refuted or significantly weakened by counterarguments see their counterpressure score decline.
Reflects the track record of the person who contributed the argument. Contributors whose past arguments have held up well over time (maintained or improved their scores) receive a modest credit. This creates an incentive for thoughtful, well-evidenced contributions. The weight is deliberately low — what matters is the argument itself, not who made it.
Evaluates whether the cited sources are genuinely independent of each other. Five news articles all citing the same original study represent one independent source, not five. Sources that arrived at similar conclusions through separate research, data collection, or analysis receive higher independence scores. This prevents citation cascades from artificially inflating evidence quantity.
Upvotes, likes, and popularity signals are displayed on every argument but are excluded from score computation. This is a hard architectural rule, not a tunable parameter. The moment popularity influences scoring, the platform becomes a consensus engine rather than a reasoning engine. ProConWiki is designed to protect the well-evidenced minority position — the argument that has strong evidence even when most people disagree. History repeatedly shows that majority opinion and evidentiary strength are different things.
A platform that claims transparent reasoning should be transparent about its own reasoning methodology. The weights above are public by design. The specific algorithms that compute each factor (how contributor credit is calculated, how counterpressure resilience is measured, how source independence is detected) are proprietary — but the framework and its priorities are fully visible.
Displayed per argument as round(evidence_strength_score × 100). For each cited evidence row:
quality = base(source_type) × recency(year) × independence × replication × sample
Base table: meta_analysis 0.95 · randomized_controlled_trial 0.90 · peer_reviewed 0.80 · government_data 0.75 · institutional_report 0.70 · data_analysis 0.65 · expert_opinion 0.55 · news_report 0.45 · industry_report 0.40 · opinion_editorial 0.25 · anecdote 0.15 · unverified 0.05. Modifiers are bounded multiplicative adjustments (recency 0.6–1.0; independence, replication and sample-size each defined per source class).
Per-argument aggregation over its evidence set, with each item weighted by argument_relevance_score: sort quality × relevance descending, take the top five, then strength = top[0] + 0.1 × Σ(s² for s in top[1:5]), clamped to [0,1]. Secondary sources contribute quality-scaled, so a weak corroborating citation adds far less than a strong one. Design intent: best-source-dominates — citation padding is structurally ineffective.
Worked example: two peer-reviewed sources whose quality × relevance land at 0.68 and 0.63 → 0.68 + 0.1 × 0.63² ≈ 0.72 → 72/100. A third source at 0.30 would add only 0.009.
reasoning_type (direct_evidence | data_analysis | expert_opinion | logical_inference) is machine-assigned display metadata describing the inferential mode; it does not currently modulate the score. Source types are machine-classified from the citation (title/venue/URL); unclassifiable sources are deliberately scored 0.05 so vocabulary violations surface rather than pass silently.
Sort a side’s distinct sources — each source_url counted once, at its highest stated strength — descending, then side_weight = top[0] + 0.1 × Σ(rest); Supports% = pro / (pro + con). Nuanced arguments carry weight in the three-way distribution shown on the claim but are excluded from the two-way Supports/Against split.
This deliberately mirrors the per-argument rule one level up: best one dominates, the rest add a diminishing 10% each. Version history. This rule replaced a plain Σ over the side on 2026-07-31, and its unit changed from arguments to distinct sources on 2026-08-22. Under a sum, argument count was a lever — splitting one argument in two raised a side’s mass without adding evidence. Under this rule that split yields 0.99 instead of 1.8, so quality decides the balance and structure does not. Unlike the per-argument score, the side weight is not clamped to 1.0 and has no top-five cutoff: it is only ever read as the ratio pro / (pro + con), where a ceiling would erase real differences rather than bound a defined range.
Epistemic scope: the balance is a statement about the quality-weighted evidence retrieved and curated for this claim, not an estimate of the full literature and not a posterior probability. Research deliberately searches both sides, but retrieval asymmetry is a real limit and we state it rather than hide it.
Applies to reasoning-graph re-scoring; the card number is the evidence-strength computation above.
Evaluates source credibility on a multi-tier hierarchy: peer-reviewed and replicated research at the top, followed by peer-reviewed but unreplicated, primary-source government or institutional data, expert analysis, quality journalism, and opinion or anecdotal sources at the bottom. Each piece of evidence cited by an argument is classified and the aggregate quality score is normalized to [0, 1]. The tier values themselves are published above — the full base table, from meta-analysis at 0.95 down to unverified at 0.05. What is proprietary is how a citation is CLASSIFIED INTO a tier from its title, venue and URL. The hierarchy is designed to reward empirical rigor and primary-source proximity.
Counts independent evidence sources cited by the argument, with diminishing returns. The first source contributes the most; additional sources provide progressively smaller marginal gains. This reflects the epistemic reality that the jump from zero sources to one is far more significant than the jump from nine to ten. The exact curve is published above rather than withheld: strength = top[0] + 0.1 × Σ(s² for s in top[1:5]) — the best source counts in full and each further source contributes with the SQUARE of its quality, so marginal value falls away quickly. What is proprietary is how each source is scored before it enters that sum, not the sum itself.
Assesses inferential validity: whether stated conclusions follow from cited premises, whether the argument addresses the claim directly (vs. strawman or tangential reasoning), and whether formal or informal logical fallacies are present. Detected fallacies reduce L(a) proportionally to severity. The fallacy detection and severity weighting methodology is proprietary.
Measures an argument's performance under adversarial scrutiny from opposing arguments. When strong counterarguments exist (high S for opposing arguments), an argument that maintains its evidence quality and logical validity scores higher on C. Arguments that have been effectively rebutted — where the counterargument's evidence directly undermines the original premises — see C decline. Untested arguments (no counterarguments exist yet) receive a neutral baseline rather than a high score. The resilience computation is proprietary.
Reflects the contributing user's historical argument quality. Computed from the contributor's past arguments' score trajectories: arguments that maintained or improved scores over time (indicating they held up as new evidence and counterarguments emerged) increase R. Arguments that degraded significantly decrease it. The weight is deliberately constrained to 10% to prevent reputation from dominating substance. The specific credit computation is proprietary.
Detects citation dependency chains and adjusts for them. If multiple cited sources trace back to the same original research, data set, or press release, they are clustered as a single effective source for quantity purposes, and the independence score is penalized. Sources that demonstrably conducted independent research, used separate data, or arrived at conclusions through different methodological approaches receive higher independence scores. The dependency detection methodology is proprietary.
Resonance signals (upvotes, likes, agreement counts) are captured and displayed as social metadata but are architecturally excluded from S(a) computation. This is enforced as a hard constraint, not a tunable weight. The exclusion is motivated by the well-documented divergence between popular agreement and evidentiary strength across domains including public health, economic policy, and scientific consensus formation. Inclusion of popularity signals would create a positive feedback loop where highly-visible arguments accumulate votes that increase their score, which increases their visibility, irrespective of their evidentiary merit. ProConWiki is a reasoning engine, not a consensus engine.
The weights and factor definitions on this page are public. The specific algorithms implementing each factor — how Q classifies source tiers, how C computes resilience under counterargument pressure, how R tracks contributor trajectories, how I detects citation dependency chains — are proprietary trade secrets. This separation is intentional: the framework and priorities are transparent, while the competitive implementation remains protected. A patent application covering the overall methodology has been filed.
The balanced synthesis displayed on each claim page is generated by AI from the scored argument set. The synthesis weighs arguments proportionally to their composite scores when summarizing the state of evidence. Arguments below a minimum score threshold are excluded from synthesis to prevent low-quality contributions from polluting the summary. The synthesis is regenerated when the argument set or scores change materially.
Some "Relevant Resources" links on claim pages are affiliate links (marked with an Affiliate chip): if you buy a book through one, ProConWiki earns a small commission from Bookshop.org at no cost to you. Every resource is hand-approved before it appears. As an Amazon Associate I earn from qualifying purchases. Affiliate relationships never influence argument ranking, evidence scoring, or synthesis — the resource system is structurally separate from the scoring pipeline and cannot feed it.
Updated 2026-08-22. The headline verdict on each claim (“Leaning no”, “Too close to call”…) is no longer computed from the percentage alone. It is extracted from the claim’s own written synthesis — the analysis that actually reads the evidence — as a structured judgment: a direction (supports / opposes / conditional), a certainty band (high / moderate / low / very low), and the named factors behind that certainty. The Supports% bar remains, computed by the transparent formula above, as a visual summary of the evidence weights — but where the bar and the written analysis disagree, the analysis governs, because a percentage cannot carry what the analysis knows (for example, that a claim holds only in a qualified form).
This mirrors the approach used in evidence-based medicine and forensic science, where decades of methodological work converged on the same two conclusions: single additive quality scores are unreliable, and evidence certainty is best expressed as a structured judgment over named factors.