
OSINTstitute
Practical intelligence training for OSINT, cyber, and AI
I have kids and I have no idea what their school is going to look like in 12 years.
Hell, nobody in edtech knows what school looks like 12 months from now. The ground keeps moving ever since LLMs started doing everyone's homework. So I stopped guessing and started building from a field that already solved the hard part: Intelligence Studies and reasoning from sources you can't fully trust.
That is intelligence work. OSINT, source verification, structured analysis. The Intelligence Community has spent decades on the exact problem the rest of us just got handed. "The inputs are unreliable, now what?"
That's the bet behind OSINTstitute Academy.
LLMs hallucinate, the web is half-bot (soon to be all bot?), screenshots are fake - verification is the new core skill. Here are the givens I am building around:
The information environment is permanently unreliable now (LLMs, bot web, fabrications). Steady state, not phase.
Humans still have to act in it. Form beliefs, make calls, eat consequences.
AI commoditizes everything that reduces to pattern-matching against existing data - including most of what current "AI literacy" courses teach.
The IC's actual contribution isn't "how to know the truth" (that's epistemics, not their lane). It's narrower and more useful: how to act on incomplete, contradictory, possibly-deceptive information without paralysis or overconfidence.
AI doesn't bear consequences for being wrong. You do. That asymmetry is the entire reason humans still need a process - not "AI literacy," but a way to make defensible calls when the inputs are unreliable.
What's already been done
Before I get too high on my own supply: calibration as a concept isn't new. LLM training research has been working on it for years (RLHF, abstention rewards, "models mostly know what they know"). Forecasting communities like Tetlock's Good Judgment Project and Metaculus have run on Brier scores for over a decade and produced measurable superforecasters. The IC has trained analysts on this stuff for a century. Decision-science academia has the textbooks. The pieces all exist.
What doesn't exist - and what I'm betting nobody else is going to ship in time - is the combination: IC tradecraft (broad verification, not just future-event forecasting) + Brier-style calibration scoring + consumer packaging + AI-saturation framing, in one product, aimed at people who aren't already analysts or quants. Model-side calibration is for the model, not the user. Forecasting communities train a narrow surface. Professional pipelines are gated. Edtech "AI literacy" courses are teaching prompt engineering and bias awareness and calling it a day. The four ingredients are sitting in four separate kitchens.
The thesis in one sentence
The durable skill of the next decade isn't prompting, tool fluency, or "media literacy" - it's running your own commit → calibrate → update loop, formalized over a century by intelligence professionals.
Three hunches I'm testing with this platform
A graduate is someone who can be wrong on purpose.
Willing to state a belief at a confidence level, knowing they might be wrong, because that's the only way feedback can improve calibration. The fence-sitter (refuses to commit, defers to AI) and the dogmatist (commits and refuses to update) are the failure modes. The graduate sits between.
Spy tradecraft transfers better to "AI literacy" than anything currently shipping under that name - Prompting is useful. Tool fluency is useful. But the durable skill is verification under uncertainty.
A course written today can't answer a question from October. The pipeline has to assume it'll be stale - The pipeline has to expect staleness. Courses need to be revisable, evidence-linked, and easy to update when the information environment changes. Quickly.
I expect to be wrong about a lot, and that is the point. I'm just one guy hoping not to get eaten by AI while hoping I only show the right thing to the dark forest. So I will adhere to the wisdom of Ward Cunningham and post the "wrong thing" to the internet then watch the right answers come in... nicely or otherwise, and use the feedback to make the product less wrong.
Platform is live at https://www.osintstitute.com/ - 12 starting courses, ~240 lessons, handful of paying customers. Small numbers right now, but a real bet costing me many hours away from those kids.
Go ahead, tell me how I am wrong. That's the point.

I’m building the OSINTstitute Academy, a hands-on training platform for intelligence analysis. One thing I keep coming back to: AI doesn’t remove human judgment from analysis. It changes the environment judgment happens inside. That matters because a lot of training still teaches tools, while the hard part is learning how your own reasoning fails under uncertainty. I wrote this as part of a community course on judgment, AI, and intelligence analysis.
Introduction
Analysis remains, at its core, a cognitive activity. Even as large-scale data processing, machine learning, and automated inference become routine features of analytic work, judgments about meaning, relevance, plausibility, and consequence are still made by people. What has changed is not the role of human cognition, but the environment in which it operates. Analysts now work within socio-technical systems in which human reasoning is intertwined with algorithmic outputs, statistical models, and automated pattern detection. Under these conditions, understanding how thinking works becomes more important, not less.
A persistent barrier to improving analysis is that individuals have limited introspective access to their own cognitive processes. People experience the outputs of thinking—conclusions, intuitions, impressions—but not the mechanisms that generate them. This limitation persists when analysis relies primarily on reading reports and synthesizing narratives, and it remains when analysis incorporates dashboards, confidence scores, and machine-generated probabilistic forecasts. What enters conscious awareness is still the endpoint of a largely invisible process.
The presence of AI systems does not alter this fundamental fact. Instead, it changes the form in which cognitive processes are stimulated and constrained. Model outputs, visualizations, and ranked recommendations become part of the analyst's perceptual field, shaping attention and interpretation before deliberate reasoning begins. As a result, many of the most consequential influences on judgment occur prior to conscious evaluation, even when analysts believe they are reasoning carefully and critically.
This revised framework integrates empirical evidence from historical intelligence failures, experimental research on human–AI interaction, and applied studies of analytic tradecraft. Rather than treating cognitive psychology as an abstract theory, it examines how cognitive constraints manifest in real analytic environments, where time pressure, organizational incentives, and technological mediation interact. The focus is not on whether analysts or machines are “better,” but on how judgment actually emerges from their interaction.
Mental Models and Bounded Rationality in Human–Machine Systems
Human reasoning operates under conditions of bounded rationality. The mind cannot process the full complexity of its environment directly. Instead, it constructs simplified internal representations—mental models—that capture what appears most salient or causally relevant. These models enable analysts to operate effectively despite constraints on attention, memory, and computational capacity. At the same time, they introduce systematic distortions by filtering information, emphasizing coherence, and suppressing ambiguity.
Mental models guide what analysts notice, how they organize information, and which explanations they find plausible. Once formed, they tend to persist, not because analysts are stubborn or careless, but because stable models reduce cognitive effort and support timely decision-making. These efficiencies come at a cost when circumstances change or when the model itself is poorly aligned with reality.
AI systems do not replace mental models. They become part of them. Analysts develop beliefs—often informal and implicit—about what a system does, how reliable it is, and what its outputs signify. These beliefs are shaped by experience, institutional norms, and interface design more than by formal evaluation. As a result, mental models of AI systems are frequently incomplete or oversimplified, which can lead analysts either to defer too readily to algorithmic outputs or to dismiss them reflexively when they conflict with established views.
Case Study: Iraq WMD and the Persistence of Mental Models
The 2003 Iraq weapons of mass destruction assessment demonstrates how mental models can trap even highly trained analysts working in well-resourced institutions. As Tracey (2007) documents, “intelligence community analysts assumed that Iraq was hiding WMD. Hence, trapped by this mindset, they narrowly pursued only one working hypothesis.”
The failure was not primarily about missing information—it was about how existing information was interpreted through a persistent mental model. Three cognitive patterns dominated:

Confirmation bias in collection and reporting.
Jervis (2006) found that “negative information [was not] solicited or reported. Agents were unlikely to press for what their sources did not observe… Negative reports rarely if ever led to requests for follow-up by headquarters whereas positive ones did.” The system was structurally biased toward confirming the prevailing hypothesis.
Failure to test alternatives.
Official investigations found no systematic use of Red Teams or Analysis of Competing Hypotheses to challenge assumptions. Alternative explanations for Iraqi behavior—such as Hussein’s need to deter Iran while eliminating actual WMD to end sanctions—were not rigorously explored.
Overconfidence despite ambiguity.
Intelligence reports displayed “excessive certainty despite ambiguous evidence” and failed to “convey explicitly to policy makers the ambiguity of their evidence” (Tracey, 2007).
The case illustrates how mental models, once established, create self-reinforcing information loops. Evidence consistent with the model is noted and weighted heavily; inconsistent evidence is dismissed or ignored. Betts (2007) characterizes this as intelligence “trying to be useful” overwhelming “being strictly accurate”—a dynamic where the desire to support decision-makers led analysts to present conclusions with greater confidence than the evidence warranted.
AI Systems and Mental Models: Empirical Evidence
The core analytic challenge, therefore, is not simply to integrate more information but to understand how simplification occurs within the human–machine system. Cognitive constraints do not disappear when machines are introduced; they are redistributed. Some simplifications are performed by algorithms, whereas others are made through human interpretation of algorithmic results. Effective analysis depends on recognizing where these simplifications occur and how they shape judgment.
Recent experimental research reveals how AI systems interact with analysts' mental models in ways that can amplify rather than mitigate bias:
Confirmation bias with AI recommendations.
Nourani et al. (2024) found that mental health professionals “were more inclined to trust and accept AI recommendations when they aligned with their initial diagnoses and professional intuition.” Crucially, “those claiming higher expertise demonstrated increased skepticism when AI’s suggestions deviated from their professional judgment.” AI outputs are accepted when they confirm existing views and discounted when they challenge them.
Stronger anchoring effects.
Burton et al. (2025), studying 775 managers, found that “the source of the recommendation (human or AI) interacted with the anchor… a high-anchor produced different performance ratings for each source.” AI recommendations carried implicit technical authority.
Limited effectiveness of cognitive interventions.
Lawless et al. (2025) found that “none of the CF interventions mitigated the influence of biased AI recommendations.” Motivation for analytical thinking mattered more than procedural interventions.
Anchoring and the Timing of Interpretation
In this environment, early interpretations take on disproportionate influence. Initial assessments, preliminary model outputs, or early alerts often anchor subsequent reasoning, especially when information is ambiguous. Once an interpretation provides a coherent account of events, the cognitive system naturally resists revision.
Worked Example:
An analyst receives an early AI alert rating a threat as “high confidence.” Subsequent reports present contradictory indicators, yet these are interpreted as noise rather than as disconfirming evidence—even after the model revises its confidence downward.
This anchoring effect is not eliminated by awareness or expertise. Confidence in one’s own judgment or in the apparent objectivity of technical systems can strengthen it. When AI-generated outputs are treated as neutral reflections of reality rather than products of specific assumptions and design choices, they can anchor thinking more powerfully than human judgments.

The contrast between the 9/11 and Iraq intelligence failures illustrates different manifestations of this dynamic. In the 9/11 case, the warning was insufficient—signals were present but not integrated. Bar-Joseph and McDermott document that “senior officers, officials, and analysts received scores of increasingly ominous warnings” that were blocked or explained away. In Iraq, excessive confidence in a flawed model led to overstated conclusions. Both failures reflect the difficulty of revising established interpretations.
Can AI Systems Reduce Cognitive Bias? A Contextualized Assessment
Understanding these dynamics requires shifting attention away from individual errors and toward systemic patterns of cognition. The question is not why analysts sometimes get things wrong, but why certain errors recur even among experienced professionals using advanced tools.
A nuanced assessment suggests AI’s impact on bias is context-dependent:
Where AI can help.
Fasolo et al. (2025) show analytics can counter selective processing, anchoring (in some contexts), and groupthink. A 2025 MDPI study demonstrates the differential effectiveness of six methodologies across five biases.
Where AI introduces new problems.
Bansal et al. (2023) document automation bias even when AI advice is erroneous. Stochl et al. (2021) identify 20 biases affecting interpretation of ML outputs.
Context-dependent effectiveness.
Success depends on task characteristics, user expertise, system design, and organizational culture.
Practical Frameworks for Human–Machine Analysis
Analysis in the AI era remains a human responsibility. Machines extend perception, memory, and computation but do not remove the need for judgment.
Structured Analytic Techniques in AI-Augmented Environments
Analysis of Competing Hypotheses (ACH)
ACH evaluates multiple explanations by testing them against evidence. In AI contexts, analyst intuitions and model outputs should both be treated as hypotheses.
Worked Example: ACH with an AI Forecast
An AI system flags increased probability of coordinated cyber intrusions.
Hypothesis A: A hostile state actor is preparing an attack.
Hypothesis B: Criminal groups are exploiting seasonal vulnerabilities.
Hypothesis C: Model artifact due to anomalous training data.
Traffic anomalies support A and B but not C. Lack of corroborating HUMINT weakens A. Recent retraining strengthens C. Provisional judgment favors B while explicitly noting assumptions about model stability.

Note: Coulthart (2017) found mixed evidence for ACH reducing confirmation bias.
Key Assumptions Check
This technique makes implicit assumptions explicit and tests their validity.
Worked Example: Assumptions in an AI Risk Score
Assumption 1: Training data reflects current conditions.
Assumption 2: Correlations imply causation.
Assumption 3: Missing data is random.
Assumption 1 is most fragile; if false, conclusions collapse. Analysts seek updated data before acting.
Red Team Analysis
Red Teaming adopts an adversary perspective.
Worked Example: Red Teaming an AI Warning
AI issues a high-confidence unrest warning. The Red Team asks how this could be wrong: media amplification inflates signals, adversaries seed misleading indicators, or the model overweights historical patterns. The exercise reveals susceptibility to information manipulation.
Organizational and System Design Principles
Recognition over correction
Structured techniques as cognitive scaffolding
Process evaluation alongside outcomes
Cultural transformation
Complementarity by design
Conclusion
The evidence converges on four conclusions:
1. Cognitive constraints persist despite expertise or technology.
2. AI transforms rather than eliminates these constraints.
3. Structured approaches work when supported by culture and design.
4. Understanding human–machine interaction remains essential.
The path forward requires sustained attention to how judgment emerges in practice within human–machine systems, rigorous evaluation of what works and why, and organizational cultures that value epistemic humility alongside technical capability. This is not primarily a technological challenge. It is a challenge of cognition, culture, and craft—made more urgent, not less, by the sophistication of the tools now available.
Continue at the OSINTstitute
I’m turning this research into a community course inside the OSINTstitute because I think the next generation of analysts needs more than tool training. They need practice noticing how judgment actually forms, especially when AI is part of the workflow.
Part II is live here: https://www.osintstitute.com/community-courses/judgement-under-uncertainty-in-human-machine-analysis
I’d genuinely value feedback from people building AI products, training platforms, cyber tools, or decision-support systems.
If you have a strong answer or a case study from your own work, I’d also be interested in publishing thoughtful community contributions on the OSINTstitute.
References
Bansal, G., et al. (2023). The effects of explanations on automation bias. ScienceDirect.
Bar-Joseph, U., & McDermott, R. (2023). Are intelligence failures still inevitable? Intelligence and National Security.
Betts, R. K. (2007). Two faces of intelligence failure: September 11 and Iraq's missing WMD. Political Science Quarterly, 122(4).
Burton, J. W., et al. (2025). How was my performance? Exploring the role of anchoring bias in AI-assisted decision making. ScienceDirect.
Coulthart, S. J. (2017). An evidence-based evaluation of 12 core structured analytic techniques. International Journal of Intelligence and Counterintelligence, 30(2), 368–391.
Fasolo, B., Heard, C., & Scopelliti, I. (2025). Mitigating cognitive bias to improve organizational decisions: An integrative review, framework, and research agenda. SAGE Journals.
Heuer, R. J., Jr., & Pherson, R. H. (2020). Structured analytic techniques for intelligence analysis (3rd ed.). CQ Press.
Jervis, R. (2006). Reports, politics, and intelligence failures: The case of Iraq. Journal of Strategic Studies, 29(1).
Lawless, E., et al. (2025). Impacts of cognitive forcing and need for cognition on biased AI-assisted decision making about mental health emergencies. Scientific Reports.
MDPI. (2025). Cognitive bias mitigation in executive decision-making: A data-driven approach integrating big data analytics, AI, and explainable systems. Electronics, 14(19).
Nourani, M., et al. (2024). Confirmation bias in AI-assisted decision-making. ScienceDirect.
RAND Corporation. (2016). Assessing the value of structured analytic techniques in the U.S. Intelligence Community (RR-1408).
Stochl, J., et al. (2021). A review of possible effects of cognitive biases on interpretation of rule-based machine learning models. ScienceDirect.
Tracey, R. S. (2007). Trapped by a mindset: The Iraq WMD intelligence failure. Air University.
UK Government. (2023). Human-centred ways of working with AI in intelligence analysis.
1 Like
Comment
About
OSINTstitute Academy is a hands-on training platform for OSINT, cyber intelligence, and AI-era source verification. It teaches practical analyst tradecraft through guided lessons, active practice, and AI-critique

4 Comments
I like the idea that the real skill is making defensible decisions under uncertainty. It reminds me of cybersecurity, where tools like "Network Threat Detection" are only effective if people know how to interpret the evidence rather than blindly trusting every alert.
It is incredibly daunting trying to figure out how to prepare the next generation for a world where the line between a verified fact and a hallucinated bot response has completely disappeared.
The real shift here is recognizing that the "Intelligence Community" model works because it focuses on the weight of evidence rather than the search for a perfect truth, which is the only way to avoid paralysis when your data source is an unreliable LLM.
Do you think that teaching people how to assign "confidence scores" to their own beliefs will be the hardest part of the curriculum for those who are used to the binary right-or-wrong nature of traditional schooling?
I'd bet it's not the hardest part, people adjust to "you can be 60% right" pretty fast once they see a score like that a few times. The sneakier, harder thing: getting people to commit to their own authority instead of having AI sit as the new "permission-to-speak" machine replacing a teacher or a textbook.
Your 'weight of evidence vs. perfect truth' line is sharper than what I wrote in the post, by the way.
Breaking the habit of using AI as a "permission-to-speak" machine is the real hurdle because once you outsource your judgment to a model, you lose the ability to defend your own conclusions.
Moving from "is this true?" to "how much evidence supports this?" is the only way to stay functional in an information environment that is permanently noisy and often deceptive.
I apply this exact principle of defensible calls in my high-tier PR and media placement work where we use verified data to build undeniable authority for brands on major news outlets.
Since the goal is to have users commit to their own authority, does the platform include a "post-mortem" feature where they can analyze why their confidence score was off after a fact is eventually verified?