Pedestrian Archetypes Extension -- More Pedestrian Models for Autonomous Vehicle Safety Testing
Seven more dangerous pedestrian types autonomous vehicles need to watch for
In earlier work, the author organized dangerous pedestrian behavior into 12 'archetypes,' but continued annotation of YouTube dash-cam footage turned up 7 more recurring behavior patterns that didn't fit the original set. This preprint defines those 7 new archetypes -- including a Con Artist who stages fake collisions, a Foreigner unfamiliar with local traffic norms, a content-creating Influencer, a Protester, an emotionally Confronted pedestrian, wheel-based Pseudo Pedestrians, and Street Vendors -- with video-frame evidence for each. Essential and optional behaviors for each archetype were derived statistically from the annotated video dataset.
METAL MEDIA explanatory visual
How the pedestrian archetype set was expanded
Evidence statusMeasured results and planned work
- Original 12 archetypes12 dangerous pedestrian types defined in the prior Pedestrian Archetypes paper, e.g. Wanderer, Drunk, Distracted
- Re-annotating dash-cam footageContinued tagging of YouTube dash-cam videos with the PedAnalyze behavior ontology, checked against existing archetypes
- New patterns identified7 recurring behavior combinations found that the original 12 archetypes could not explain
- Essential/optional behaviors derived40%+ of examples = essential behavior, 10-39% = optional behavior, set statistically from the data
- 7 new archetypes definedCon Artist, Foreigner, Influencer, Protester, Confronted, Pseudo Pedestrian, and Street Vendor, each with video evidence
What they did
- Continuing to annotate pedestrian-vehicle risk footage from YouTube dash-cams, the author found 7 recurring behavior patterns not explained by the original 12 archetypes (Wanderer, Drunk, Distracted, Flash, and others).
- The classification followed four steps: annotate behaviors using the PedAnalyze behavior ontology, compare against the original 12 archetypes to check for genuinely new patterns, select representative video clips and extract key frames, and derive essential/optional behavior tags from the annotated data.
- A behavior was labeled essential if it appeared in at least 40% of an archetype's examples, and optional if it appeared in 10-39%, giving a data-driven way to separate core from associated behaviors.
- The 7 new archetypes are the Con Artist (stages collisions for insurance fraud), the Foreigner (misreads traffic due to unfamiliar norms), the Influencer (occupies roads to film content), the Protester, the Confronted (hostile/emotional reactions), the Pseudo Pedestrian (moves fast on wheels like skateboards or wheelchairs), and the Street Vendor.
- Each archetype is illustrated with real dash-cam scenes (a staged accident, a confused tourist at a 3-way Venice intersection, a Christmas video shoot in New York traffic, a truck deliberately hitting a protester, etc.) and summarized in Tables I-VII of essential/optional behaviors.
![Fig. 1: Pedestrian from [2] stages an accident with friend.](https://media.metallab.ai/papers/2607.16922/f0.png)
| Essential Behaviors | Optional Behaviors |
|---|---|
| • Collision • Looking • Fall • Not-Cross • Run into traffic • Cross without crosswalk | • Climbing onto carhood • Thrown-back • Ignore traffic |

| Essential Behaviors | Optional Behaviors |
|---|---|
| • Ignore traffic • Near-miss • Cross without crosswalk • Run into traffic • Not looking/glancing | • Retreat • Back-turned • Collision |

| Essential Behaviors | Optional Behaviors |
|---|---|
| • Ignore traffic • Gesturing • Glancing |

| Essential Behaviors | Optional Behaviors |
|---|---|
| • Frozen • Not-cross • Looking • Ignore-traffic • Agitated | • Group-walk • Along-lane • Back-turned |
![Fig. 2: Confused Venice tourist unsure what to do. [4]](https://media.metallab.ai/papers/2607.16922/f4.png)
| Essential Behaviors | Optional Behaviors |
|---|---|
| • Agitated • Aggression • Cross • Ignore traffic | • Assault • Gesturing • Near-miss |

| Essential Behaviors | Optional Behaviors |
|---|---|
| • Ignore traffic • Run into traffic • Cross • Collision • Not looking/glancing | • Along-lane • Swerve • Pop-out-occlusion |
![Fig. 3: Influencer with dog as photographer lies down [7].](https://media.metallab.ai/papers/2607.16922/f6.png)
| Essential Behaviors | Optional Behaviors |
|---|---|
| • Pause-start • Pickup-object • Cross-on-red |
![Fig. 4: A group of pedestrians from [9] blocks the road as a vehicle passes through. One pedestrian is hit and remains in the path, causing a second collision.](https://media.metallab.ai/papers/2607.16922/f7.png)
Findings
- Continued annotation of YouTube dash-cam videos revealed 7 recurring pedestrian behavior patterns not captured by the original 12 archetypes.
- Essential behaviors (appearing in 40%+ of examples) and optional behaviors (10-39%) were derived for each new archetype from the annotated dataset.
- Each new archetype was illustrated with a real dash-cam scene, such as a staged fraud collision, a confused tourist, a roadside video shoot, a truck hitting a protester, an agitated mother, a skateboarder's fall, and street vendors approaching cars.

Where it can be used
- Could inform the design of rare and unpredictable pedestrian scenarios for autonomous vehicle safety simulations.
- Could serve as a higher-level classification layer on top of individual behavior tags in pedestrian annotation work.
- Could be used as a framework for collecting representative examples of risky pedestrian-vehicle interactions by archetype.

Limits and open work
- The study relies solely on YouTube dash-cam videos, without quantitative discussion of sample size, geographic, or cultural bias.
- The exact sample sizes underlying the 40% essential / 10-39% optional thresholds are not specified in the text.
- No results are yet reported showing these new archetypes actually applied in an autonomous vehicle safety test or simulation.
- This is a preprint and has not undergone peer review as a finalized publication.

Why it matters
Autonomous vehicle safety testing needs to prepare not just for rule-following pedestrians but for rare, unpredictable, and risky behaviors, and this work systematically names and defines such behaviors so they can be used in annotation, simulation, and evaluation. Grouping behaviors into archetypes fills a gap that single behavior tags leave open -- explaining why a pedestrian retreats or crosses unpredictably, not just what they did.
![Fig. 5: A hostile mother [3] as she crosses with children.](https://media.metallab.ai/papers/2607.16922/f11.png)
Terms in this paper
- Pedestrian Archetype · A collection of behaviors that together identify a specific type of pedestrian
- PedAnalyze · The author's prior pedestrian behavior ontology and standardized tagging framework
- Essential/Optional Behavior · Essential = appears in 40%+ of an archetype's examples; optional = appears in 10-39%
- Dash-cam footage · Video recorded by a vehicle-mounted camera, the primary data source for this study
Original abstract (English)
In our prior work, Pedestrian Archetypes, we defined pedestrian archetypes as collections of behaviors that uniquely identify a specific type of pedestrian. The first paper proposed 12 pedestrian archetypes, including the Wanderer, Drunk, Distracted, Flash, Indecisive, Blind, Flock, Jaywalker, Elderly, Kid, Eventful, and Parked Pedestrian. These archetypes were introduced to move beyond single behavior labels and provide a more natural way to describe how dangerous pedestrians actually behave pr
Read on arXivLatest papers
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?AI coding agents were tested on fixing real scientific software, and even the best one failed more than half the time
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM ServingMaking sparse attention fast enough and accurate enough for real LLM serving, not just papers
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM AgentsMaking customer-service AI agents follow the whole procedure, not just avoid one bad action
- EXIMO: VLM Guided Exploration of VLA PoliciesTeaching a robot new chores without human teleoperation, by letting a chatty AI supervise it
- EnvHarness: Awakening Static Worlds for Agent LearningInstead of building new training worlds from scratch, this work adds a plug-in layer that reshapes existing ones around each agent's actual weaknesses
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the ModelCompanies that rent AI instead of owning it can only do half of AI safety oversight
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressA fix for AI models that get penalized by their teacher even when they're reasoning correctly
- PersonalBench: Measuring the Authorship Gap in LLM PersonalizationAI can be prompted to write 'like someone,' but its own voice never fully disappears
Latest from METAL MEDIA
Figures: Taorui Huang et al., arXiv:2607.16922, cc-by-nc-sa-4.0