Abstract
Generative artificial intelligence can produce linguistically clear, structured, and persuasive answers whose claims are nevertheless false, incomplete, unsupported, or unsuitable for the relevant context. The central problem is not merely the well-known possibility of so-called hallucinations. It lies in the growing gap between linguistic plausibility and epistemic reliability.
The article uses the pointed term bullshit detector neither as a psychological diagnosis of the machine nor as a supposedly reliable human alarm mechanism. What is meant is a personal, embodied, risk-based review process that examines statements, presentation, use, and institutional responsibility separately. The aim is neither blanket distrust nor blind acceptance but calibrated trust. One could also say: use common sense.
At the same time, the article shows why individual critical faculties are not enough. The less users are able to conduct expert scrutiny and the greater the potential harm, the less the safety of an AI system may depend on their personal judgment. Providers, organizations, and public institutions bear an upstream responsibility. The bullshit detector is therefore not merely an individual skill. It is a social and institutional infrastructure for dealing with artificial intelligence. But of course, it is also no excuse not to continue developing one’s own sixth sense.
1. The Problem Is Not Nonsense
People have always produced nonsense. They make mistakes, exaggerate, confuse correlation with causation, cite incorrectly, oversimplify complex relationships, and speak with great confidence about things they have not adequately examined.
Artificial intelligence did not invent this problem, but it has changed the form of nonsense.
A person with weak expertise reveals their uncertainty through hesitation, contradictions, limited vocabulary, or recognizable gaps in knowledge. Generative AI does not display these signals. It can formulate an unsupported statement with the same linguistic quality as a well-substantiated one. A real study and an invented study can appear in the same citation style. A conjecture can sound like the state of research.
Generative AI does not therefore necessarily reduce the difference between truth and error. But it can reduce the visible difference among knowledge, conjecture, and invention.
That is precisely why we need a bullshit detector. And it is not a machine; it is that small part of the brain that has existed since time immemorial. It is simply not yet very well trained for AI.
The term does not mean a program that detects whether a text was written by AI. The origin of a text says little about its quality. A human-written text can be false and insubstantial. An AI-assisted text can be excellently researched, properly sourced, and intellectually independent.
So does a statement deserve to be believed, shared, or made the basis of a decision?
2. What “Bullshit” Means Here
Harry Frankfurt distinguished bullshit from lying. The liar orients himself toward the truth because he wants to conceal it. The bullshitter behaves differently. What matters is not whether his statement happens to be true or false. What matters is his lack of commitment to the question of whether it is true. [1]
Michael Townsen Hicks, James Humphries, and Joe Slater applied this analysis to ChatGPT. They argue that the linguistic production of large language models is not oriented toward truth in the same way as a reasoned human factual claim. [2]
A language model has no human attitude. It is not honest, dishonest, or indifferent. It pursues no conscious intent to deceive. “Bullshit” is therefore neither a psychological diagnosis of the machine nor an established technical term for an error.
Here, the term describes a constellation:
Guiding PrincipleThe linguistic persuasiveness of a statement exceeds its epistemic support.
An answer sounds competent without being adequately supported. An explanation appears plausible even though it does not reliably reflect the actual process by which it was produced. A list of sources looks scientific even though some sources do not exist or fail to support the relevant claims.
3. Not an Inner Warning Light but a Review Process
The term “detector” is misleading. It sounds as though people could intuitively and reliably recognize when an AI answer does not hold up.
They cannot. But they can train themselves. And even if it sounds a little absurd, the more they work with AI, the more developed their own detector becomes. (Source: Personal anecdotal experience)
A layperson often cannot recognize a false medical recommendation, an invented legal principle, or a flawed statistical interpretation from its style. Persuasively formulated errors can be especially overwhelming for nonexperts.
The bullshit detector is therefore not an inner warning light. It is a risk-based review process that examines four levels.
3.1 The Output Level
Is the specific statement correct, supported, and complete enough for the context?
This concerns facts, figures, sources, definitions, conclusions, and omitted qualifications.
3.2 The Presentation Level
Is the epistemic status of the statement communicated appropriately?
Is something certain, probable, disputed, hypothetical, or unknown? Does the system make these differences visible, or does it smooth them over linguistically?
3.3 The Use Level
Does the person rely on the answer appropriately?
A good answer can be ignored. A bad one can be adopted without scrutiny. Not only model performance but also the relationship between the person and the system determines the outcome.
3.4 The Institutional Level
Who reviews, assumes responsibility for, and corrects the output?
An AI-generated sentence in a private birthday card has a different status from an AI-assisted medical recommendation, an administrative decision, or a scientific publication.
An invented source is an output error. Concealed uncertainty is a presentation error. Uncritical adoption is a use error. Missing control processes are an institutional error.
The bullshit detector must recognize all four, but not by the same means.
4. Not Every “AI” Is the Same
A more precise technical distinction is also required.
4.1 The Base Model
A base model generates outputs on the basis of statistically learned patterns. Linguistic probability does not guarantee factual correctness.
4.2 The Application System
A modern AI product can additionally use web search, databases, retrieval systems, calculators, code execution, source retrieval, and automated verification mechanisms.
Such a system has a different epistemic status from a freely generating base model. Reliable data sources and verification steps can significantly reduce errors.
A search function can select an unreliable source. A retrieval system can overlook the relevant passage. A correctly retrieved source can be misinterpreted. A citation can exist without supporting the claim.
4.3 Institutional Use
The third level is the concrete process of use.
Who uses the system? For what task? With what data? Under what time pressure? What review is planned? Who may object? Who is liable?
The same technology can be unproblematic in private brainstorming and highly risky in a clinical or legal procedure.
So the question is not: How good is the model? No, the question is: How reliable is the entire system in this specific use case?
5. Why Plausibility Is So Persuasive
People derive a large share of their knowledge from the statements of others. We cannot make every observation ourselves, replicate every study, or independently derive every law. Societies therefore depend on reasoned trust.
Dan Sperber and his coauthors describe the human mechanisms for evaluating communicated information as epistemic vigilance. Among other things, people assess a source’s competence and trustworthiness as well as a message’s plausibility and consistency. [3]
Generative AI encounters this vigilance under unusual conditions.
It displays characteristics that we often interpret as signals of trustworthiness in people:
linguistic fluency, rapid responses, orderly argumentation, apparent neutrality, a broad vocabulary, calm authority.
In a person, these characteristics may be connected with competence. In AI, they are initially properties of the generated surface.
Fluency does not prove understanding. Confidence does not prove certainty. Balance does not prove completeness. A citation does not prove that the source supports the claim.
Pennycook and colleagues investigated human receptiveness to pseudo-profound statements. Their findings show that people do attribute depth to formulations that sound meaningful but are weak in substance. [4]
Generative AI can lower the cost of formulating and adapting plausible statements in large quantities. The long-term effects this will have on judgment and public communication have not yet been conclusively researched.
The linguistic quality of an answer can no longer be regarded as a reliable indication of its factual quality.
6. Hallucination Is Only One Class of Error
Research documents that generative systems can produce statements unsupported by sources, inputs, or facts. Review articles distinguish among different forms of such hallucinations and discuss causes and countermeasures. [5]
Unfortunately, focusing on completely fabricated statements does not go far enough, because a broad field lies between a correct statement and a complete invention:
a correct figure with the wrong reference year, a real study with an overstated finding, a minority opinion presented as consensus, an accurate rule without a crucial exception, a correlation reframed as causation, a plausible summary that removes the central context, a source that exists but does not substantiate the claim.
For the models examined at the time, TruthfulQA showed that widespread human misconceptions can be reproduced and that larger models were not automatically more truthful on this benchmark. [6] A study of generative search engines found substantial gaps in the complete support of claims and in the accuracy of individual citations. [7]
These findings relate to specific models, systems, and points in time. They must not be generalized as an immutable property of every future AI system.
Guiding PrincipleA professionally formulated answer may require substantial verification even if it contains no conspicuous hallucination.
7. The Goal Is Calibrated Trust
These risks do not provide a reasonable argument for blanket distrust.
AI can perform certain tasks more accurately, quickly, and consistently than people. Anyone who categorically rejects good machine recommendations also creates errors.
Research on automation therefore distinguishes among appropriate use, overreliance, and unjustified nonuse. Lee and See speak of appropriate reliance, meaning reliance aligned with the task and the system’s performance. [8]
The goal is calibrated trust.
People should neither follow an AI because it is an AI nor contradict it for that reason.
Which tasks is the system strong at? Which errors typically occur? How current is its information? Which data sources does it use? How great would the harm from an error be? Can the answer be independently verified?
This is far more demanding than skepticism. Skepticism can be indiscriminate. Calibrated trust requires discernment.
8. Explanations Are Not an Automatic Antidote
A widespread hope is this: If AI explains its answer, people can assess its quality more effectively.
That can work. But it does not have to.
Chen and colleagues investigated how people use different types of explanations in AI-assisted decisions. In their tasks, feature-based explanations sometimes led to greater overreliance, while example-based explanations provided better conditions for meaningful correction. [9]
In five studies involving a total of 731 participants, Vasconcelos and colleagues showed that people are more likely to examine explanations when the expected benefit justifies the cognitive cost. Difficult tasks, explanations that are hard to verify, and a lack of incentives can promote overreliance. [10]
An explanation is therefore not good merely because it sounds explanatory.
People must be able to understand, examine, and compare the explanation with alternatives. Under time pressure or without subject-matter expertise, a plausible explanation can make a false recommendation even more persuasive.
The bullshit detector within us must therefore also ask:
Does this explanation enable control, or does it merely simulate traceability?
9. The Case in Favor: AI Can Demonstrably Help People
A serious analysis must not treat the benefits as an aside. Because that is simply what one does in analyses.
In a controlled experiment, Noy and Zhang found that participants completed certain typical professional writing tasks more quickly with generative AI and achieved results that received higher average ratings. [11]
Brynjolfsson, Li, and Raymond studied the use of an AI assistant by more than 5.000 customer-service employees. In this specific use case, access to assistance increased productivity, measured by cases resolved per hour. Less experienced employees benefited particularly strongly. [12]:
AI can make knowledge more accessible, reduce language barriers, accelerate routine tasks, provide guidance to beginners, and distribute experiential knowledge within organizations.
They do not show that AI improves every form of knowledge work. The findings are specific to the task, system, and context.
A CHI study involving 319 knowledge workers found associations between greater trust in generative AI and lower self-reported critical effort. Greater confidence in one’s own competence, by contrast, was associated with more intensive scrutiny. The study shows a possible shift in cognitive work toward monitoring and integrating AI outputs. It does not prove a general or long-term decline in human thinking ability. [13]
Guiding PrincipleAI can improve productivity and access to knowledge. Whether this produces greater competence, dependency, or loss of competence depends on the task, user, system design, and organization.
10. The Competence Paradox
Less experienced people can benefit particularly strongly from AI. At the same time, they are more likely to lack the expertise needed to recognize subtle errors.
The beginner receives a professional-looking answer more quickly.
The expert is more likely to recognize where it does not hold up.
This creates a competence paradox:
Those who receive the greatest support may also become the most dependent on it.
This is not an argument against AI-assisted learning. AI can explain extremely well, answer follow-up questions, generate examples, and provide individualized feedback.
What matters is the form of use.
Someone who asks:
“Explain the path to the solution, identify uncertainties, and then test my understanding,”
uses AI differently from someone who asks:
“Solve the task and write the submission.”
In the first case, AI can support the development of competence. In the second, it can improve the visible result without building the underlying ability.
The bullshit detector therefore examines more than the machine.
It also examines one’s own division of labor with it.
11. The Societal Consequences
11.1 Education
Education is not merely about producing correct results. It develops abilities: reading, comparing, reasoning, formulating, doubting, failing, and correcting.
AI can support these processes. But it can also shortcut them.
Schools and universities should therefore teach more than prompting. Learners must dissect AI answers, check sources, recognize overstated claims, develop counterpositions, and explain content without AI.
The relevant object of assessment is shifting.
Whether someone can produce a polished text is becoming less important. Whether they understand, defend, and correct it is becoming more important. So it is really no different from school in the past, when we copied our homework from the top student.
11.2 The World of Work
In many knowledge processes, the cost of generation can fall more quickly than the cost of expert review.
A report can be generated in a few minutes. Its figures, sources, and conclusions must still be checked. If these review costs are passed on to colleagues or customers, individual speed may increase while the productivity of the overall system declines.
Organizations must therefore measure more than output. They must take account of review effort, the cost of errors, and competence development.
11.3 Science
Scientific quality does not arise from an academic tone. It arises from methods, data, citation, criticism, and reproducibility.
A source must exist. It must have been read. It must support the specific claim. A figure must be traceable to the original. The authors remain responsible for every sentence.
“The AI named the source” is not a scientific justification.
11.4 The Public Sphere
Generative AI lowers the cost of linguistically professional content. As a result, the volume of plausible texts can grow faster than society’s capacity to scrutinize them.
When verified research and synthetic plausibility look the same, epistemic exhaustion threatens: the attitude that it is no longer possible to determine what is true anyway.
11.5 Inequality
AI can democratize access to knowledge. At the same time, there is a risk of a new inequality between people who can assess AI outputs with subject-matter expertise and people who have to rely on their surface appearance.
Without public education and verification infrastructures, the ability to assess information can become a new driver of inequality.
The advice “Check the answer” protects no one who has neither the time nor the expertise nor access to reliable sources.
12. Responsibility Does Not Begin with the User
It would be convenient to derive just one demand from all of this:
People need to think more critically.
That would allow providers and organizations to pass the risk on to individual users.
The order of responsibility must begin the other way around.
First: providers
Providers must design systems so that known errors are reduced, sources are traceable, and uncertainties are visible. They must not systematically present linguistic certainty as stronger than the evidence warrants.
Second: organizations
Organizations must specify which tasks AI may be used for, which checks are required, who is responsible, and when a process must be stopped.
Third: public institutions
Legislation and oversight must safeguard audits, liability, complaints, education, and access to independent information.
Fourth: users
Users need epistemic vigilance. But they must not be the final safeguard for systems they are not professionally equipped to control.
Key ideaThe less opportunity individuals have to verify something and the greater the potential harm, the less reliability may depend on a personal bullshit detector.
Article 4 of the EU AI Act requires providers and deployers to take measures to ensure a sufficient level of AI literacy among their staff and other persons acting on their behalf.
But AI literacy must not be confused with tool proficiency or prompt training. It must also encompass the ability to interpret outputs appropriately, recognize limitations, and stop when the evidence is insufficient.
13. The Conflict with Product Logic
A system that frequently indicates uncertainty, makes counterpositions visible, directs users to primary sources, and does not answer when the evidence is insufficient may feel less seamless than a system that immediately provides a clear answer.
Products are often optimized for speed, frequency of use, low abandonment rates, and deep integration. Epistemic quality, by contrast, may require friction:
Time, verification, follow-up questions, not knowing, contradiction, sometimes no answer.
That does not mean companies inevitably build bad systems. It means that truth quality does not automatically follow from a good user experience.
This is a question of system design, business models, and institutional countervailing power.
14. A Practical Verification Process
A useful bullshit detector does not always apply the same level of rigor. It begins with the risk.
Step 1: Determine the consequences
What happens if the answer is wrong?
Nonbinding brainstorming requires less scrutiny than a medical, legal, financial, or public factual claim.
Step 2: Separate the claims
Which statements are facts? Which are interpretations? Which are forecasts? Which are recommendations? Which are value judgments?
Only separate claims can be properly verified.
Step 3: Clarify the system’s status
Is a base model answering from the patterns learned during training? Is the application using current web sources? Is it accessing a defined database? Was a calculation performed, or was the answer merely estimated in words?
Step 4: Check the sources
Does the source exist? Are the author, title, and date correct? Is it a primary source? Does it support the specific claim? Has its finding been overstated?
Step 5: Deconstruct the numbers
Numbers need a reference quantity, time period, population, method, and comparison value.
“40 percent” is not reliable information until it is clear: 40 percent of what, when, among whom, and compared with what?
Step 6: Look for counterpositions
Which credible perspective contradicts the answer? What limitation would a subject-matter expert mention? What information could change the conclusion?
Step 7: Triangulate independently
Key statements need an independent means of verification: a primary source, specialist database, statutory text, official statistics, systematic review, or qualified expert.
Multiple answers from related AI systems are not automatically independent evidence.
Step 8: Be able to stop
If a statement cannot be adequately verified, the correct label is:
Not verified.
That is not failure. It is epistemic hygiene.
15. The Counterarguments
People produce bullshit too
Correct. AI is not the original cause.
But it can change the production costs, speed, and adaptability of plausible content. This can turn a human problem into a problem of scale.
Constant verification destroys the utility
Also correct, if every harmless task is treated like a high-risk process.
That is why verification must be risk-based. The greater the consequences, the stricter the standard.
AI can make better decisions than people
For certain tasks, yes.
The bullshit detector is not meant to make human authority absolute. It is meant to safeguard against human and machine errors by setting them against each other.
Sometimes a person must correct the AI. Sometimes the AI must correct the person.
Good systems enable both.
No one can verify everything themselves
That is precisely why societies depend on the division of labor and justified trust.
The bullshit detector does not mean personally recalculating everything. It means recognizing when ordinary trust is sufficient and when a source, process, or expert is needed.
The term is too polemical
It would be unsuitable for a technical standard. In public debate, it serves a purpose: it starkly identifies the difference between persuasive language and unearned epistemic authority.
For the term to hold up scientifically, its limits must remain clear.
It does not describe machine consciousness.
It describes a verification problem for society.
16. Conclusion: The Machine May Answer
We will use AI.
This morning. For emails, research, translations, analyses, teaching, programming, administration, science, and decisions.
The decisive skill no longer lies solely in generating good answers.
It lies in being able to assess the status of an answer.
Is this knowledge? A plausible assumption? An incomplete summary? A genuine source? A recommendation without sufficient basis? Or a linguistically impressive form that does not hold up epistemically?
The bullshit detector does not protect against every error.
It protects the distinction between language and knowledge, between explanation and evidence, between assistance and authority.
The more reliable they become, the less often errors may occur. But rare errors may become harder to detect because high reliability can weaken active oversight. Automation research has long warned against the misuse, disuse, and overuse of technical systems. [15]
The machine may answer.
But no person, company, or institution may regard a convincing answer as knowledge solely because it is convincing.
References
Key ideaFrankfurt, H. G. On Bullshit. Princeton University Press, 2005. Originally published as an essay in Raritan Quarterly Review, 6(2), 81–100, 1986.
Key ideaHicks, M. T., Humphries, J., and Slater, J. “ChatGPT is bullshit.” Ethics and Information Technology, 2024.
Key ideaSperber, D., Clément, F., Heintz, C., Mascaro, O., Mercier, H., Origgi, G., and Wilson, D. “Epistemic Vigilance.” Mind & Language, 25(4), 359–393, 2010. DOI: 10.1111/j.1468-0017.2010.01394.x.
Key ideaPennycook, G., Cheyne, J. A., Barr, N., Koehler, D. J., and Fugelsang, J. A. “On the Reception and Detection of Pseudo-Profound Bullshit.” Judgment and Decision Making, 10(6), 549–563, 2015.
Key ideaJi, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y., Chen, D., Dai, W., Chan, H. S., Madotto, A., and Fung, P. “Survey of Hallucination in Natural Language Generation.” ACM Computing Surveys, 55(12), 1–38, 2023. DOI: 10.1145/3571730.
Key ideaLin, S., Hilton, J., and Evans, O. “TruthfulQA: Measuring How Models Mimic Human Falsehoods.” Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, 3214–3252, 2022. DOI: 10.18653/v1/2022.acl-long.229.
Key ideaLiu, N. F., Zhang, T., and Liang, P. “Evaluating Verifiability in Generative Search Engines.” Findings of the Association for Computational Linguistics: EMNLP 2023, 7001–7028, 2023. DOI: 10.18653/v1/2023.findings-emnlp.467.
Key ideaLee, J. D., and See, K. A. “Trust in Automation: Designing for Appropriate Reliance.” Human Factors, 46(1), 50–80, 2004. DOI: 10.1518/hfes.46.1.50_30392.
Key ideaChen, V., Liao, Q. V., Vaughan, J. W., and Bansal, G. “Understanding the Role of Human Intuition on Reliance in Human-AI Decision-Making with Explanations.” arXiv:2301.07255, 2023.
Key ideaVasconcelos, H., Jörke, M., Grunde-McLaughlin, M., Gerstenberg, T., Bernstein, M. S., and Krishna, R. “Explanations Can Reduce Overreliance on AI Systems During Decision-Making.” arXiv:2212.06823, 2022.
Key ideaNoy, S., and Zhang, W. “Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence.” Science, 381(6654), 187–192, 2023. DOI: 10.1126/science.adh2586.
Key ideaBrynjolfsson, E., Li, D., and Raymond, L. R. “Generative AI at Work.” NBER Working Paper No. 31161, National Bureau of Economic Research, 2023; revised version.
Key ideaLee, H.-P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., and Wilson, N. “The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects from a Survey of Knowledge Workers.” Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, Article 1121, 1–22, 2025. DOI: 10.1145/3706598.3713778.
Key ideaEuropean Union. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence. Article 4: AI literacy. Official Journal of the European Union, 2024.
Key ideaParasuraman, R., and Riley, V. “Humans and Automation: Use, Misuse, Disuse, Abuse.” Human Factors, 39(2), 230–253, 1997. DOI: 10.1518/001872097778543886.
Key ideaLong, D., and Magerko, B. “What Is AI Literacy? Competencies and Design Considerations.” Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems, 1–16, 2020. DOI: 10.1145/3313831.3376727.
Key ideaLaupichler, M. C., Aster, A., Schirch, J., and Raupach, T. “Artificial Intelligence Literacy in Higher and Adult Education: A Scoping Literature Review.” Computers and Education: Artificial Intelligence, 3, 100101, 2022. DOI: 10.1016/j.caeai.2022.100101.
Key ideaBergstrom, C. T., and West, J. D. Calling Bullshit: The Art of Skepticism in a Data-Driven World. Random House, 2020.
Key ideaOreskes, N. Why Trust Science? Princeton University Press, 2019.