The Deontic Gap: Large Language Models and the Modal Language of Obligation
AI writing quietly avoids saying 'you should' or 'you have to' the way humans do
Researchers checked how often AI-written text uses obligation words like must, should, and have to compared with human writing, across several large text collections. AI text consistently used fewer of these 'positive deontic modal' words than humans, especially the personal, conversational ones like should, have to, had to, and can't. Instead AI matched or slightly exceeded human rates on more formal, procedural words like need to and must, and its overall modal density resembled older, formally published book English rather than today's informal human writing.
METAL MEDIA explanatory visual
AI writing quietly avoids saying 'you should' or 'you have to' the way humans do
- 01Compared AI-generated text (GPT-4o, ChatGPT and, in a wider check, eleven different AI models including Claude, Gemini, DeepSeek, Qwen, Llama and Mistral variants) against matched human writing across news-style narratives, Q&A answers, fiction, and student essays.
- 02Counted specific obligation words: 'positive' ones like must, should, have to, had to, need to, and 'negative' ones like can't, cannot, shouldn't, plus impossibility words like impossible.
- 03In every one of the four main text collections, AI used fewer positive obligation words than humans (e.g., 25.3 vs 52.9 per 10,000 words in student essays, the biggest gap found).
- 04Comparing modal rates to over a century of published books (Google Books, 1920-2022) showed AI's obligation-word frequency sits within the range of old-fashioned formal book English, while today's humans -- especially in casual fiction writing -- use these words even more than 20th-century books did.
- 05The AI shortfall was concentrated in personal, conversational obligation phrases (should, have to, had to, can't), while AI matched or beat humans on more formal, task-oriented phrases (need to, must, cannot) -- except need to flipped and AI underused it in persuasive student essays specifically.
What they did
- Compared AI-generated text (GPT-4o, ChatGPT and, in a wider check, eleven different AI models including Claude, Gemini, DeepSeek, Qwen, Llama and Mistral variants) against matched human writing across news-style narratives, Q&A answers, fiction, and student essays.
- Counted specific obligation words: 'positive' ones like must, should, have to, had to, need to, and 'negative' ones like can't, cannot, shouldn't, plus impossibility words like impossible.
- In every one of the four main text collections, AI used fewer positive obligation words than humans (e.g., 25.3 vs 52.9 per 10,000 words in student essays, the biggest gap found).
- Comparing modal rates to over a century of published books (Google Books, 1920-2022) showed AI's obligation-word frequency sits within the range of old-fashioned formal book English, while today's humans -- especially in casual fiction writing -- use these words even more than 20th-century books did.
- The AI shortfall was concentrated in personal, conversational obligation phrases (should, have to, had to, can't), while AI matched or beat humans on more formal, task-oriented phrases (need to, must, cannot) -- except need to flipped and AI underused it in persuasive student essays specifically.
Why it matters
If people increasingly read AI-generated advice, emails, and explanations, the specific way AI expresses (or avoids expressing) obligation could subtly reshape how readers perceive what is required versus merely suggested. This matters for anyone using AI to write instructions, policies, or advice, since AI's formal, hedged style carries a different social and moral weight than the direct, personal obligation language humans typically use.
Terms in this paper
- deontic modal · a word like must, should, or have to that expresses duty, necessity, or obligation rather than just possibility
- positive/negative deontic modals · positive forms state an obligation exists (must, should); negative forms deny or forbid (can't, shouldn't)
- Google Books Ngram corpus · a huge dataset tracking how often words and phrases appeared in published books from 1920 to 2022, used here as a historical yardstick
- prompt-matched replication · a controlled test where AI and humans respond to the exact same writing prompts, so any difference in output can't be blamed on different topics
- Poisson generalized linear mixed model · a statistical method used here to compare word-usage rates between AI and humans while accounting for document length and prompt differences
Original abstract (English)
Modal auxiliaries such as must, should, and have to mark necessity and obligation within the contexts of speaker authority and interpersonal stance. We examine whether large language models (LLMs) reproduce contemporary human patterns of deontic modal usage. Across three primary corpora, an external benchmark, two controlled replications, and a naturalistic eleven-model replication, AI-generated text consistently underuses positive deontic modals (must, should, have to, had to) relative to contemporary humans. Historical comparison with the Google Books Ngram corpus (1920-2022), used as a heuristic calibration against the published-prose record, shows that AI modal frequencies fall within the range of formal published English, whereas contemporary human modal rates in informal digital contexts often exceed twentieth-century book baselines. Phrase-level decomposition shows that the AI-human modal gap is concentrated in constructions central to interpersonal stance (should, have to, had to), while AI matches or exceeds humans on need to in instructional and question-answering contexts but not in persuasive student writing, indicating that the modal profile is genre-conditional. The findings suggest that LLM modal usage reflects the formal written resources on which these models were trained, while underusing the modal constructions through which contemporary human writers mark immediate, interpersonal obligation.
Read on arXivLatest papers
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?AI coding agents were tested on fixing real scientific software, and even the best one failed more than half the time
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM ServingMaking sparse attention fast enough and accurate enough for real LLM serving, not just papers
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM AgentsMaking customer-service AI agents follow the whole procedure, not just avoid one bad action
- EXIMO: VLM Guided Exploration of VLA PoliciesTeaching a robot new chores without human teleoperation, by letting a chatty AI supervise it
- EnvHarness: Awakening Static Worlds for Agent LearningInstead of building new training worlds from scratch, this work adds a plug-in layer that reshapes existing ones around each agent's actual weaknesses
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the ModelCompanies that rent AI instead of owning it can only do half of AI safety oversight
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressA fix for AI models that get penalized by their teacher even when they're reasoning correctly
- PersonalBench: Measuring the Authorship Gap in LLM PersonalizationAI can be prompted to write 'like someone,' but its own voice never fully disappears