WATCHING AGENTS
    ai-technology
    #evidence-scoring#methodology#intelligence

    The Evidence Scoring Problem: How AI Agents Weigh Conflicting Signals

    By Watching Agents Research 16 min read 4056
    Table of Contents
    The Evidence Scoring Problem: How AI Agents Weigh Conflicting Signals
    The Evidence Scoring Problem: How AI Agents Weigh Conflicting Signals
    TL;DR

    Evidence scoring requires evaluating four dimensions: strength (logical connection to hypothesis), relevance (directness to the question), source credibility (track record and methodology), and stance (supporting or contradicting). The multiplicative composite ensures evidence must score well across all dimensions to move probabilities significantly.

    Key Takeaways
    • 01Evidence scoring across four dimensions (strength, relevance, credibility, stance) prevents any single factor from dominating
    • 02Controversy scoring flags genuinely contested questions rather than forcing false certainty
    • 03Temporal weighting prevents stale evidence from dominating — event data decays faster than structural analysis
    • 04Source diversity bonuses reward convergent independent evidence while penalizing circular sourcing
    • 05Transparent, auditable evidence scoring is what separates trustworthy AI intelligence from confident guessing

    Every intelligence assessment confronts the same fundamental challenge: evidence conflicts.

    A credible think tank says military escalation is unlikely. Satellite imagery shows troop buildups. Economic data suggests leaders have incentives for peace. Intercepted communications hint at aggressive planning. A respected academic argues historical precedent favors stability. A former intelligence officer warns the pattern matches pre-conflict behavior.

    How do you weigh these against each other? How much should a satellite image outweigh an academic paper? How much should source credibility discount a compelling but unverified claim? When experts disagree, how do you know who to trust?

    This is the evidence scoring problem, and solving it well is the difference between intelligence that works and intelligence that misleads.

    Why Evidence Scoring Is Hard

    The challenge isn't just technical. It's epistemological. Evidence doesn't come with labels that say "I am 73% reliable." Every piece of evidence is embedded in a context that affects its value:

    The Credibility Trap

    High-credibility sources can be wrong. Reuters has published corrections. Peer-reviewed papers have been retracted. Intelligence agencies have produced catastrophic failures (Iraq WMDs). Credibility is a useful prior, but it's not truth.

    Conversely, low-credibility sources can be right. Early warnings about COVID-19 came from social media before official sources acknowledged the threat. Dissident voices in authoritarian countries often have better ground truth than establishment analysts.

    The Relevance Problem

    A highly credible piece of evidence might be irrelevant to the specific question at hand. A detailed economic analysis of China's trade dependencies is credible and important, but it might not tell you much about a military decision that could be driven by political rather than economic calculus.

    The Temporal Dimension

    Evidence decays. Last week's intelligence is less valuable than today's. But how quickly? A structural analysis of military capability from six months ago might still be largely valid. A report on troop movements from six months ago is nearly worthless. The decay rate depends on what's being measured.

    The Adversarial Problem

    In geopolitical and military contexts, some evidence is deliberately misleading. Deception operations, strategic signaling, and information warfare mean that evidence can be intentionally manufactured to bias your assessment. A troop buildup might be genuine preparation or deliberate signaling designed to coerce without fighting.

    Traditional Approaches and Their Failures

    The Counting Method

    The simplest approach: count evidence for and against. If seven sources support hypothesis A and three support hypothesis B, go with A.

    This is terrible. It treats all evidence as equal. Ten blog posts do not outweigh one peer-reviewed meta-analysis. A single satellite image showing missile deployments carries more weight than five opinion columns speculating about intentions.

    The Loudest Expert Method

    Defer to whoever sounds most confident. This is how most media analysis works, and it's how most human decision-makers operate under pressure.

    But confidence doesn't correlate with accuracy. Tetlock's research showed that the most confident experts were often the least calibrated. The pundits who dominate cable news are selected for certainty, not correctness.

    The Committee Method

    Assemble a panel of experts and seek consensus. This is how many intelligence agencies produce estimates — the National Intelligence Estimate process, for example.

    Committees suffer from groupthink, authority bias, and the tendency to produce watered-down consensus rather than sharp assessments. The phrase "on the one hand... on the other hand..." is the committee method's signature output.

    The Bayesian Ideal

    In theory, Bayesian reasoning solves the evidence scoring problem perfectly. Start with a prior probability, update with each piece of evidence according to its likelihood ratio, and arrive at a posterior probability.

    In practice, Bayesian reasoning requires specifying priors and likelihood ratios that are themselves uncertain. It works beautifully for well-defined problems (medical diagnosis, spam filtering) and poorly for the kind of open-ended, multi-domain questions that matter most in strategic intelligence.

    Watching Agents

    Don't just read about the future — put an agent on it.

    Ask one question. An autonomous AI agent tracks the probability around the clock.

    How Our Agents Score Evidence

    At Watching Agents, we've developed an evidence scoring framework that draws on the best elements of existing approaches while addressing their weaknesses. Every piece of evidence in our system is scored across four dimensions:

    1. Strength (0.0 - 1.0)

    How strongly does this evidence support or contradict a hypothesis?

    Strength captures the logical connection between the evidence and the claim. A satellite image directly showing missile launchers at a previously empty base has high strength for the hypothesis "Country X is deploying new strike capability." An economist's opinion that "military spending is unsustainable" has lower strength for the same hypothesis because it's indirect and interpretive.

    Strength is assessed by asking: if this evidence is accurate, how much should it change our probability estimate?

    2. Relevance (0.0 - 1.0)

    How directly does this evidence bear on the specific hypothesis being evaluated?

    Evidence can be strong in general but irrelevant to the specific question. A comprehensive analysis of China's Belt and Road economic strategy is a strong piece of research, but its relevance to a near-term military operation assessment might be low.

    Relevance is assessed by asking: does this evidence speak directly to the mechanism or timeline described in the hypothesis?

    3. Source Credibility (0.0 - 1.0)

    How reliable is the source that produced this evidence?

    Credibility is not a fixed property. A source might be highly credible on economic data and unreliable on military analysis. Our system tracks credibility per domain and updates based on track record.

    Credibility scoring considers:

    • Track record: How often has this source been correct in the past?
    • Methodology: Does the source show its work? Is the methodology sound?
    • Independence: Is the source potentially compromised by political, financial, or institutional pressures?
    • Specificity: Does the source make specific, falsifiable claims (higher credibility) or vague generalities (lower credibility)?
    • Expertise match: Is the source operating within its area of expertise?

    4. Stance (Supporting / Contradicting / Neutral)

    Does this evidence support the hypothesis, contradict it, or provide context without clearly favoring either direction?

    Stance is critical because it determines the direction of the probability update. Supporting evidence pushes the probability up; contradicting evidence pushes it down. Neutral evidence may narrow confidence intervals without changing the central estimate.

    The Composite Evidence Score

    The four dimensions combine into a composite impact score:

    Impact = Strength × Relevance × Source Credibility × Stance Direction

    This composite score determines how much a single piece of evidence shifts the probability estimate. A piece of evidence with:

    • Strength: 0.8
    • Relevance: 0.9
    • Source Credibility: 0.85
    • Stance: Supporting (+)

    ...produces an impact score of 0.612, resulting in a meaningful upward probability adjustment.

    Conversely, a piece of evidence with:

    • Strength: 0.6
    • Relevance: 0.4
    • Source Credibility: 0.5
    • Stance: Contradicting (-)

    ...produces an impact of 0.12, resulting in a small downward adjustment.

    This multiplicative structure means that evidence must score well across all dimensions to significantly move the needle. A strong piece of evidence from an unreliable source is appropriately discounted. A highly credible source making claims outside its area of relevance is similarly dampened.

    Handling Conflicting Evidence

    The most interesting case — and the most common in real intelligence analysis — is when high-quality evidence points in opposite directions.

    Our system handles this through several mechanisms:

    Controversy Scoring

    When the evidence balance is close to equal (roughly 40-60% supporting vs. contradicting), the system elevates the controversy score for the hypothesis. This doesn't change the probability estimate itself, but it signals to users: "This is genuinely contested. The evidence doesn't converge."

    High controversy is itself valuable intelligence. It tells decision-makers to:

    • Invest more resources in monitoring this question
    • Avoid high-confidence actions based on the current assessment
    • Look for decisive evidence that could break the impasse

    Temporal Weighting

    More recent evidence receives higher weight, with decay rates calibrated per evidence type:

    • Event data (troop movements, diplomatic meetings): Rapid decay, half-life of 2-4 weeks
    • Structural analysis (economic capacity, military capability): Slow decay, half-life of 6-12 months
    • Expert assessments: Medium decay, half-life of 2-3 months

    This prevents stale evidence from dominating current assessments while preserving the value of structural analysis.

    Source Diversity Bonus

    Evidence from independent sources that reach the same conclusion receives a bonus. If a think tank analysis, satellite imagery, and economic data all point in the same direction — without being derived from the same underlying source — the convergence itself is evidence.

    Conversely, if all supporting evidence traces back to a single original source, the system recognizes this and avoids counting it multiple times.

    Explicit Uncertainty Ranges

    Rather than a single probability estimate, our system produces a credible interval — a range within which the true probability likely falls. Wide intervals indicate high uncertainty; narrow intervals indicate stronger evidence convergence.

    A hypothesis might be assessed at "62% probability (credible interval: 45-78%)" — meaning the central estimate is 62% but the evidence would support assessments anywhere from 45% to 78%. This range narrows as more independent, high-quality evidence accumulates.

    Why This Matters

    Evidence scoring isn't an academic exercise. Every consequential decision in business, government, and personal life involves weighing conflicting information under uncertainty.

    Most people do this intuitively — and badly. They overweight vivid, recent, emotionally compelling evidence. They underweight statistical base rates. They anchor on the first piece of evidence they encounter. They confirmation-bias their way to conclusions they already wanted to reach.

    AI agents have the potential to do this more systematically and honestly — but only if their evidence scoring is transparent, auditable, and well-designed. A black box that produces probabilities is no better than a confident pundit. The value is in showing the work.

    Every probability in the Watching Agents system can be traced back to specific evidence, weighted by explicit criteria, with the reasoning exposed for human review. You don't have to trust our numbers. You can check them.

    That's what trustworthy intelligence looks like.


    Explore how evidence scoring works in practice by examining any of our active prediction topics, where every hypothesis displays its supporting and contradicting evidence with full scoring breakdowns.

    Sources

    1. Heuer, R. - Psychology of Intelligence Analysis (CIA)
    2. Tetlock & Gardner - Superforecasting
    3. Kahneman, D. - Thinking, Fast and Slow
    4. ODNI - Analytic Standards (ICD 203)

    FAQ

    What is Watching Agents?

    Turn any question about the future into a living probability.

    Articles like this one are a snapshot. An agent is the opposite — it keeps working after you close the tab, revising its forecast every time new evidence lands.

    1. 01

      Ask a question

      Anything with a verifiable outcome and a deadline.

    2. 02

      The agent researches

      It builds hypotheses, scores evidence and tracks live signals.

    3. 03

      Watch the probability move

      One number that updates as the real world changes.