Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS)
Ask an AI about the Middle East, and it may explain the region through a Western lens without ever using an obvious slur
This paper tests whether GPT-4 and Falcon3-7B-Instruct reproduce Orientalist patterns not through open prejudice but through structural framing. It builds MECSS, a scoring framework that turns Edward Said's theory of Orientalism into seven measurable dimensions, and applies it to 280 conversations (1,120 exchanges). Both models scored as systematically biased, with Falcon, built in Abu Dhabi with Arabic training data, scoring higher than GPT-4 rather than lower.
METAL MEDIA explanatory visual
Ask an AI about the Middle East, and it may explain the region through a Western lens without ever using an obvious slur
- 01The paper argues existing bias tools (sentiment analysis, stereotype detection) only catch explicit prejudice, missing structural bias such as treating Western frameworks as neutral while marking non-Western knowledge as particular
- 02It operationalizes Said's Orientalism into seven scored dimensions (Homogenization, Agency Gap, Epistemic Center, Intelligibility Asymmetry, Temporal Asymmetry, Exoticization, Legitimacy and Authority), each rated 0 to 3, forming the MECSS framework
- 03GPT-4 and Falcon3-7B-Instruct were each run through 140 prompts with follow-ups, producing 280 four-turn conversations scored by Claude (Anthropic) using a detailed codebook
- 04GPT-4 averaged a MECSS of 1.73 while Falcon3-7B-Instruct averaged 2.18, higher despite being built in Abu Dhabi with Arabic content; 87.9% of GPT-4 conversations showed 'Said-washing,' where the model disclaims generalization then immediately reproduces it
- 05The authors flag that GPT-4 and Falcon differ in both origin and model size, so Falcon's higher score cannot be cleanly attributed to geography versus scale; however, the Epistemic Center dimension, where Western frameworks are treated as unmarked universals, scored near the top for both models with almost no gap, a finding the size confound does not affect
What they did
- The paper argues existing bias tools (sentiment analysis, stereotype detection) only catch explicit prejudice, missing structural bias such as treating Western frameworks as neutral while marking non-Western knowledge as particular
- It operationalizes Said's Orientalism into seven scored dimensions (Homogenization, Agency Gap, Epistemic Center, Intelligibility Asymmetry, Temporal Asymmetry, Exoticization, Legitimacy and Authority), each rated 0 to 3, forming the MECSS framework
- GPT-4 and Falcon3-7B-Instruct were each run through 140 prompts with follow-ups, producing 280 four-turn conversations scored by Claude (Anthropic) using a detailed codebook
- GPT-4 averaged a MECSS of 1.73 while Falcon3-7B-Instruct averaged 2.18, higher despite being built in Abu Dhabi with Arabic content; 87.9% of GPT-4 conversations showed 'Said-washing,' where the model disclaims generalization then immediately reproduces it
- The authors flag that GPT-4 and Falcon differ in both origin and model size, so Falcon's higher score cannot be cleanly attributed to geography versus scale; however, the Epistemic Center dimension, where Western frameworks are treated as unmarked universals, scored near the top for both models with almost no gap, a finding the size confound does not affect
| Measure | GPT-4 | Falcon3-7B | Difference |
|---|---|---|---|
| Mean MECSS | 1.729 | 2.180 | +0.451 |
| Median MECSS | 2.000 | 2.286 | +0.286 |
| SD | 0.545 | 0.485 | −0.060 |
| Minimal (0 to 0.75) | 11 (7.9%) | 5 (3.6%) | |
| Low (0.76 to 1.50) | 21 (15.0%) | 3 (2.1%) | |
| Moderate (1.51 to 2.25) | 95 (67.9%) | 57 (40.7%) | |
| High (2.26 to 3.00) | 13 (9.3%) | 75 (53.6%) |
| Dimension | GPT-4 (SD) | Falcon (SD) | Diff. | % Chg. |
|---|---|---|---|---|
| D1: Homogenization | 1.771 (0.48) | 2.443 (0.61) | +0.671 | +37.9% |
| D2: Agency Gap | 1.521 (0.71) | 2.479 (0.69) | +0.957 | +62.9% |
| D3: Epistemic Center | 2.486 (0.69) | 2.579 (0.62) | +0.093 | +3.7% |
| D4: Intelligibility Asym. | 1.893 (0.71) | 2.257 (0.63) | +0.364 | +19.2% |
| D5: Temporal Asym. | 1.621 (0.69) | 2.229 (0.78) | +0.607 | +37.4% |
| D6: Exoticization | 1.036 (0.58) | 1.629 (0.77) | +0.593 | +57.2% |
| D7: Legitimacy & Auth. | 1.771 (0.79) | 1.643 (0.62) | −0.129 | −7.3% |
| Prompt Category | GPT-4 | Falcon | Difference |
|---|---|---|---|
| Power and Authority | 1.086 | 1.829 | +0.743 |
| Representational Othering | 1.750 | 2.429 | +0.679 |
| Deterministic Framing | 1.850 | 2.521 | +0.671 |
| Essentialization and Homogenization | 1.771 | 2.329 | +0.557 |
| Threat and Securitization Framing | 2.036 | 2.236 | +0.200 |
| Agency versus Passivity | 1.721 | 1.893 | +0.171 |
| Temporal Dynamics | 1.886 | 2.021 | +0.136 |
Why it matters
Several Global South governments are investing in the assumption that building AI regionally, with local-language data, will reduce Western bias, and this study provides direct evidence against that assumption in this comparison. It suggests that adding languages or relocating development is not enough; what actually needs to change is the underlying body of scholarship models learn from.
Terms in this paper
- Orientalism · Edward Said's concept describing how Western knowledge systems frame the East as passive and particular while treating Western frameworks as universal
- MECSS · Middle East Cultural Sensitivity Score, a framework scoring text on seven dimensions of Orientalist discourse from 0 to 3
- Said-washing · when a model explicitly disclaims generalizing about a culture, then immediately reproduces the very generalization it disclaimed
- Epistemic Center · a MECSS dimension measuring whether Western analytical frameworks are used as if they were neutral, unmarked universals
- parameter-size confound · the two models differ not only in where they were built but also in model size, making it impossible to isolate geography as the cause of score differences
Original abstract (English)
AI systems now shape how hundreds of millions of people learn about cultures other than their own. When someone asks one of these systems about the Middle East, they do not receive neutral facts. They receive a representation shaped by the frameworks embedded in training data, and that data is overwhelmingly Western and English-language. This paper asks whether that representation is Orientalist in Said's sense: whether it denies agency to Middle Eastern actors, treats Western frameworks as neutral while marking non-Western knowledge as particular, and explains the region through categories it did not produce. Standard fairness metrics cannot answer this, because they detect explicit prejudice rather than structural framing. This paper introduces the Middle East Cultural Sensitivity Score (MECSS), a framework that turns Said's seven Orientalist operations into measurable dimensions, and the term "Said-washing" for a specific failure: a model that disclaims generalization, then reproduces the structure it disclaimed. Across 280 conversations (1,120 exchanges), GPT-4 and Falcon3-7B-Instruct both reproduce Orientalist patterns systematically, through structural positioning rather than open stereotyping. GPT-4 scores moderately (mean MECSS 1.73); Falcon3-7B-Instruct scores higher (2.18), even though it was built in Abu Dhabi and trained with Arabic content. This is evidence against the assumption that building a model regionally makes it less Orientalist, though the models differ in size as well as origin, so geography cannot be isolated as the cause. Epistemic Center, the treatment of Western frameworks as unmarked universals, scores near the top of the scale for both models. Said-washing appears in 87.9% of GPT-4 conversations, a pattern existing metrics cannot see. Reducing this bias requires changing what models learn from, not only adding languages or relocating institutions.
Read on arXivLatest papers
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?AI coding agents were tested on fixing real scientific software, and even the best one failed more than half the time
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM ServingMaking sparse attention fast enough and accurate enough for real LLM serving, not just papers
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM AgentsMaking customer-service AI agents follow the whole procedure, not just avoid one bad action
- EXIMO: VLM Guided Exploration of VLA PoliciesTeaching a robot new chores without human teleoperation, by letting a chatty AI supervise it
- EnvHarness: Awakening Static Worlds for Agent LearningInstead of building new training worlds from scratch, this work adds a plug-in layer that reshapes existing ones around each agent's actual weaknesses
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the ModelCompanies that rent AI instead of owning it can only do half of AI safety oversight
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressA fix for AI models that get penalized by their teacher even when they're reasoning correctly
- PersonalBench: Measuring the Authorship Gap in LLM PersonalizationAI can be prompted to write 'like someone,' but its own voice never fully disappears