AI Papers, Explained
Five papers a day from arXiv, read and explained — not a dump of everything published.
- SWE-bench Science: Can Coding Agents Resolve Engineering Tasks in Science?AI coding agents were tested on fixing real scientific software, and even the best one failed more than half the time
- FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM ServingMaking sparse attention fast enough and accurate enough for real LLM serving, not just papers
- PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM AgentsMaking customer-service AI agents follow the whole procedure, not just avoid one bad action
- EXIMO: VLM Guided Exploration of VLA PoliciesTeaching a robot new chores without human teleoperation, by letting a chatty AI supervise it
- EnvHarness: Awakening Static Worlds for Agent LearningInstead of building new training worlds from scratch, this work adds a plug-in layer that reshapes existing ones around each agent's actual weaknesses
- From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAGEdge devices running RAG need to decide on the fly how much to compress retrieved text, or they waste energy
- Scientific Data Skills: Enabling Agent-Ready Scientific Data Services at ScalePackaging scientific datasets as ready-to-use 'skill' files so AI agents can find and understand them on their own
- SafeBranch: Branch-Pair Safety Alignment for Embodied AgentsTeaching a robot AI to know exactly which past decision was unsafe, by rewinding the sim to that moment and showing both choices side by side
- Write Once, Run Everywhere: The Axon DSL for Shape-Safe and Framework-Agnostic LLM ArchitecturesA new language, Axon, lets you write an LLM once and run it on PyTorch, JAX, MLX or vLLM
- A Virtual Member of a Community of Practice for the Society of Petroleum Engineers: From Prototype to DeploymentAn AI colleague for oil and gas engineers, ATHENA, goes from prototype to real deployment
- Hallucination as a Feature, not a Defect: Evaluating a multi-agent architecture to transform speculative language-model outputs into testable scientific hypothesesAn experiment that keeps AI 'making things up' but forces it into a controlled loop to produce testable research ideas
- Represented but Ignored: A Causal Account of Prosodic Underuse in Audio-Language ModelsAudio AI models can correctly hear tone of voice but still fail to use it when answering
- NepOOC-M: Bilingual Nepali-English Benchmark and Comparative Analysis of Multimodal Architectures for OOC DetectionNepali fake-news detector matches image+text model using text alone
- Time-Series Retrieval for Grounding Multimodal Language Models in Remaining Useful LifeShowing an AI similar past breakdown cases makes its machine-failure predictions better
- Can Conversational AI loosen Us-Versus-Them Boundaries? The Effects of Common, Dual, and Separate Identity Framings on Pro-Immigrant Intergroup HelpingFive chats with an AI chatbot made White Americans see Latine immigrants more as fellow Americans, and more willing to help them
- Linguistic Holonomy and Statistical Watermarks: Inner Geometry of Meaning-Preserving TransformationsFor AI text watermarks, where you edit matters more than how much you edit
- Remember, Verify, or Ask? Cross-Family Evaluation of Memory Commitment in LLM AgentsAI assistants would rather double-check facts than ask you a question, even when asking is the right call
- When Machines Speak: A Unified Generative Framework for Integrating Machine-Native Symbols into Pretrained Large Language ModelsA framework that lets pretrained LLMs generate machine-only codes -- like recommendation item IDs or legal citation markers -- right alongside plain text
- LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive AlignmentPeeking at a few early training gradients before fine-tuning starts to set up LoRA smarter
- Learning how to Forget: Fine-tuning for Long-Context Sparse AttentionTeaching AI models to forget the right things when reading very long documents
- Active Inference as Context Acquisition for AI AgentsTeaching AI agents to decide when asking a question is worth the tokens it costs
- Interaction valence reveals contrasting social networks in dairy cattleWhen you split cow interactions into 'friendly' vs 'fighting', the herd's social map looks completely different
- Automated Summarization of Financial News Using Large Language Models and Retrieval-Augmented Generation: An Early Empirical Study (Fall 2023)Testing AI summaries of stock news, the simple approach beat the trendy retrieval-based one
- Can Agent Memory Systems Track Evolving State?AI agents keep answering with facts that were already updated, and this paper measures that failure directly
- Rethinking the Evaluation and Optimization of LLM-Based Social SimulationWhen AI mimics human survey answers, checking if it 'got the answer right' is the wrong way to grade it
- Beyond Imitation: Filtering On-Policy Distillation by Reasoning ProgressA fix for AI models that get penalized by their teacher even when they're reasoning correctly
- Bounded Sovereignty and the Control Tax: Pricing AI Oversight When the Deployer Does Not Own the ModelCompanies that rent AI instead of owning it can only do half of AI safety oversight
- PersonalBench: Measuring the Authorship Gap in LLM PersonalizationAI can be prompted to write 'like someone,' but its own voice never fully disappears
- ADAPT: Physics-Aware Diffusion-based World Models for Adaptive Predictive Transferable HVAC ControlAn AI that keeps buildings comfortable and efficient even when the season or climate changes, by teaching it the physics of heat
- GenMatch: An End-to-End Generative Matching Framework for Micro-View Order-Dispatching in Ride-HailingDiDi replaced its multi-step ride-hailing dispatch pipeline with one generative model and saw real-world gains
- Causal Reasoning with Bipartite Graphical Causal ModelsA bathtub example exposes a blind spot in standard causal reasoning tools
- How to Navigate Uncertainty About AI ConsciousnessWe may never prove whether an AI is conscious, but we can check if it has states that would feel good or bad
- Spike-based Belief Propagation in Nonlinear Dynamical SystemsA brain-like spiking neural network learns to park a car using probabilistic reasoning
- Rethinking Patch Based Multivariate Time Series Forecasting with Semantic Structured PartitioningA new way to slice time series into meaningful chunks instead of arbitrary equal-length pieces
- A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware EvaluationBuilding speech recognition for Mizo, a low-resource Indian language, by fine-tuning Whisper and SraVaani on 17.62 hours of new data
- Forking Fast: Efficiently Estimating Uncertainty Dynamics in Text GenerationA statistical smoothing trick lets you find where an LLM's reasoning could branch into different answers without resampling hundreds of times at every step
- SynFlow: A Multidimensional Diachronic Semantic Analysis ToolkitAn open-source tool that breaks down how a word's meaning changed, not just that it changed
- LLMs as Acquisition Policies for Finite-Pool Materials Optimization: A Controlled StudyResearchers tested whether LLMs can pick which new material to try next, instead of classic statistical methods
- Robust Incomplete Multimodal Sentiment Analysis via Iterative Proxy CorrectionWhen text input is missing or broken, this AI doesn't guess once and move on—it revises its guess step by step to read emotions more reliably
- Natural Language Code Retrieval for 1C:Enterprise: An Open Benchmark and Efficient Bi-EncoderA first-of-its-kind search benchmark and AI model let you find 1C business-software code using Russian-language questions
- A knowledge-guided agentic framework for mitigating patient-context ambiguity in health queriesGetting AI to ask the right follow-up questions before answering vague health queries
- Stopping and Routing LLM Judge PanelsA method for deciding how many AI judges to call, and when to stop calling more
- The Asymmetric Harms of LLM CompressionShrinking AI models can quietly erode common knowledge more than rare facts, while keeping models confidently wrong and hiding bias shifts inside stable-looking averages
- A Strong Linear Baseline for Whole-Heart Cardiac Shape Completion on CT, with an Open Eleven-Structure Statistical Shape ModelFor filling in missing heart structures on CT scans, a simple math formula beat a graph deep-learning model
- Reliable Financial Named Entity Recognition under Domain ShiftAn AI's confidence trained on formal filings turns unreliable once it reads tweets
- Robust Metaheuristics under Uncertainty for Berth Allocation and Quay Crane Assignment: A ReviewA review of algorithms that keep port ship-and-crane schedules from falling apart when things go wrong
- Generating Diverse Personas for User Simulators to Test Interview Dialogue SystemsTo test interview-style chatbots you need many different fake users, so this work has an LLM automatically generate those fake user personalities
- Are LLMs becoming similarly creative? Evidence from three years of modelsNewer AI chatbots are giving increasingly similar answers to each other, three years of data show
- ReguSim: Evaluating LLM Agent Rule Grounding in Financial ComplianceAI trading agents say they know the rules, but still place orders that break them
- Asymmetric Attention Heads: Structured Head-Wise Context Allocation for Transformer AttentionTesting whether every attention head in a Transformer really needs to see the same amount of context
- Automatic bioinformatic software named entity recognition from literatureA new AI tool automatically spots software and database names buried in biology papers
- Auditing Cross-Lingual Fairness in Language Model WatermarkingAI text watermarks that are supposed to catch machine-written content work far less reliably in many non-English languages, and the gap tracks language families, not individual languages
- Bringing analytic rigor to agentic AI for science: The Brain Researcher platform for neuroimaging data analysisA system that makes AI show its work when analyzing brain-imaging data, not just deliver an answer
- Outcome Monitors: Recovery Affordances for Silent Tool FailuresA no-force warning note that just says 'here's what else you can try' cuts AI agents' silent-failure blind spot in half
- Transformer Models for Text Summarization: A Comparative Study of BART, BERT, and RoBERTaHead-to-head test of three summarization AIs shows BART, which rewrites text from scratch, beats models that just pick existing sentences
- When Saying No Makes Better Videos: Designing Dual Gatekeeping for Pedagogically Grounded AI Content CreationBetter teaching videos come from AI systems that know when to say no
- TESTNAV: Pareto-Guided Search for Compositional Robustness TestingA smarter way to test AI models against combined real-world glitches, without checking every possible combination
- One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business WorkflowsSucceeding once doesn't mean an AI agent can be trusted - a 507-task benchmark exposes the reliability gap
- ReCache: Efficient KV Cache Reuse and Compression for Tool-Augmented LLM AgentsMaking AI agents stop re-reading the same tool manuals every time, for a 3x-plus speed-up
- When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language ModelsSlipping an irrelevant sentence into a prompt shifts multimodal AI answers in a predictable, formula-like way
- Compliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System MessagesForcing image-understanding AI to follow hidden system rules quietly wrecks its accuracy, and it collapses even more when users push back
- SAPO: Single-Rollout Autoregressive Policy Optimization for Agentic Reinforcement LearningA way to train AI agents better using just one attempt per task, with no separate evaluator model needed
- Specification-delta-driven data governance: an empirical study of the «spec-delta» as the unit of change in lakehouse data platformsShould data-platform changes be reviewed as a 'spec-delta' instead of a code diff? A study design, not results yet
- Symposium: Trust via Auditable Records for Communities of AI Scientist AgentsA record-keeping system that stops AI research agents from quietly faking or hiding how they got their results
- Learning Early-to-Final Solution Consistency for MILP AccelerationTeaching AI to judge which early guesses in optimization problems will stay correct, to speed up solvers
- Towards general embodied intelligence: integrating large language models, knowledge bases, and reasoning capabilities to build the next generation of AI agentsA survey on what it takes to turn talkative AI into robots that can actually act in the world
- Hear2Act: Benchmarking When Prosody Should Change What an Assistant DoesA 480-scenario benchmark tests whether tone of voice actually changes what an AI assistant decides to do
- TT-net: Quantum Inspired Tensor Network Denoising in Conditional GANsLetting image channels talk to each other improves GAN-based denoising
- Beyond Memory Majority: Latent-Source Reasoning for Multi-Agent Memory ArbitrationWhen multiple AI agents' memories are secretly copies of the same source, this method stops the system from being fooled by a fake majority
- Frequency-Aware Continual Learning for Smart Contract Vulnerability Detection with Large Language ModelsA smart-contract vulnerability detector that keeps learning new bug types without forgetting old ones, then folds everything into one model
- Learning Hierarchical Skill Policies with Offline Quality-Diversity Reinforcement LearningTeaching robots to pick out only the useful moves from messy, mixed-quality recorded data
- Enforcing LLM Safety through DMD-based Classification of Prompt-Response Embedding DynamicsJudging whether an AI's answer is 'safe' by watching how its words move through math space, not by reading them like a human
- Air Traffic Control Using Large Language Models: Prompt Engineering, Architecture, and EvaluationResearchers put AI in the air traffic controller's seat and found it sounds right but often gives the wrong instructions
- Mitigating Identity Essentialism in LLM Agents with Longitudinal Life TrajectoriesGiving AI characters a life story, not just a demographic label, stops them from acting like clones of their group
- Projector Is All You TrainYou don't need to retrain the whole language model to teach it a new sense like 3D
- Interrupting the Loop: Periodic Subject Changes Raise Judged Surprise and Connection in Base Language ModelsLetting a base AI model ramble alone, but yanking it onto a new subject every few hundred words, makes its output feel more surprising and more connected
- Dynamic Gated Cross-Modal Fusion with Sarcastic-aware Contrastive Regularization for Multimodal Sarcasm DetectionWhen text and image seem to agree but secretly clash, that mismatch is sarcasm — and this AI learns to spot it
- HealMed: Multilingual Evaluation of Large Language Models in MedicineMedical AI chatbots ace English but stumble badly in Swahili and Zulu
- ExPhy: A Benchmark for Explicit Physical Property Learning in Multi-Object Trajectory ForecastingPredicting where objects go doesn't mean a model actually knows their weight, friction, or bounciness
- Optimal Skill Selection for LLM Agents with Provable Bicriteria GuaranteesA method that picks which 'skill documents' to feed an AI coding agent, with mathematically guaranteed near-optimal results
- FM-Bench: A Benchmark for Long-Horizon Management with Competing AgentsHanding an AI a football club to run for 20 years reveals that winning comes from management habits, not raw model power
- Improving Natural-Language Combinatorial-Optimization Accuracy in Resource-Constrained Language Models via Formal AbstractionsA small language for scheduling problems lets weaker AI models write feasible schedules instead of broken code
- The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation ExplanationsNetflix treats its AI grader for recommendation blurbs as a living system, not a one-time build
- Optimized Fuzzy Logic Approach with the IEEE Key Gas Method for Diagnosing Power Transformer Faults Using Dissolved Gas AnalysisA fuzzy-logic upgrade to transformer gas diagnosis that separates out carbon monoxide to cut misdiagnoses
- Are LLMs Safe Beyond Text: Do Emojis Expose Gaps in Safety EvaluationSlipping past AI safety filters with emojis instead of words
- FinSkillBench: Evaluating AI Agents and Domain Skills for Investment ManagementGiving AI investment agents a validated playbook boosts performance sharply, but letting them write their own doesn't help
- The Deontic Gap: Large Language Models and the Modal Language of ObligationAI writing quietly avoids saying 'you should' or 'you have to' the way humans do
- On the Triangle Inequality for the Jaccard Distance in Arbitrary LatticesThe math behind Jaccard distance, the classic similarity score, now proven to work as a true distance on much broader mathematical structures than sets
- RDFdL: Integrating RDF with Differential Dynamic LogicA framework lets you ask a single graph query that checks both 'is this physical transition safe' and 'who is the technician responsible for it'
- StocksTalk: A Voice-Enabled Conversational Agent for Structured Query Generation over Web DataA voice assistant that turns spoken stock-screening requests into checked SQL queries
- Stability-Aware Feature Design for Robust Watermark Detection in Machine-Generated TextA watermark-detection method that survives repeated paraphrasing of AI-written text
- DeepTCM1.0: A Multi-Expert AI Agent for Deciphering Mechanisms of Chinese Herbal Formulae Based on General Large Language ModelsA team of 11 AI expert agents built to explain how a classical Chinese herbal formula actually works
- Accurate Decoding of Natural Sentences from Non-Invasive Brain RecordingsAn AI that reads full sentences from brain scans, cutting word errors to 39%
- Fractional Decay KV-Cache: Ownership-Aware Memory Management for Improved Inference Relevancy in Dialog SystemsKeeping a chatbot's memory from clinging to yesterday's topic
- Redakto - The Incognito Tab for LLMsAn open-source tool that strips names and addresses before text hits an LLM, without hurting downstream accuracy
- MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned EvaluatorsA framework and a small AI model that check whether images and text respect big social values like peace, justice, and freedom
- Self- and Other-Labels Induce Bidirectional Bias in LLM JudgesAI judges change their scores based on who they're told the author is, not on actual quality
- Latent Space Refusal Anchoring for Low-Resource African Languages: Mechanistic Safety Recovery Without RetrainingAI models that refuse harmful requests in English but comply in African languages get fixed by tweaking internal signals, without any retraining
- Governance Records as Supervision: Verifier-Selected Self-Training for Structured Workflow RepairLetting an outside verifier pick which AI-generated plans are correct can turn a cheap model into a reliable one
- UMER: Unifying Embedding and Ranking via Pair-Aware Discriminative Reasoning for Universal Multimodal RetrievalA universal search AI learns to compare candidates side by side to explain why one is the right match
- Temporal Multi-Signal Fusion for Token-Level Hallucination DetectionCatching AI's made-up facts works better when you read them as a stretch of text, not word by word
- Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking SchemeGrading AI exams the wrong way inflates their scores: a benchmark built on Vietnam's real high-school marking rules
- Emergence of Agentic AI: A Review on Evolution, Background, Working Principles, Applications, Adoption Factors, and Future Research DirectionsA review paper maps out how AI evolved from rule-based systems to today's goal-chasing, self-directed Agentic AI
- Language Models for Portuguese: A Systematic Mapping StudyA full map of 46 Portuguese-language AI models, built for the first time
- When Clean Signals Are Not Enough: Detecting Structural Ambiguity for Safe Wearable Stress ClassificationA stress-detecting wearable AI averaged 93% accuracy but totally failed on one person, so researchers built a pre-check that flags risky readings before the AI even makes a guess
- Backdoor Learning in Language Models and Vision-Language ModelsA PhD thesis maps out how 'backdoors' can be planted in and detected from language models and vision-language models, using the attention mechanism as the key clue
- BERTilda: Explainable Topic Lifecycle Tracking with Split/Merge Detection via Similarity-and-Flow Temporal GraphsA method that tracks how topics split, merge, and disappear over time, and shows its reasoning
- Position: Multi-Agent Systems Should Prioritize Concurrency ControlMulti-agent AI systems fail not because agents can't 'coordinate', but because they're not built to handle concurrent edits to shared data
- WhiteMatter: All-to-All Cross-Layer Connections via KV MixingLetting every Transformer layer read every other layer's memory, like white-matter wiring in the brain, boosts language-model quality without adding depth
- Which Negatives Matter? Ask Your Text Encoder: Adaptive Similarity Margins for Dense-Caption RetrievalAn image-to-long-caption search AI kept thinking it had already solved the problem, so it never learned to tell near-identical captions apart
- When Do LLMs Actually Help? Evaluating LLMs as Data Quality AnnotatorsLLMs aren't a universal fix for data cleaning: their value depends heavily on the task
- Figurative and Cultural Knowledge in LLMs: Investigating Cross-Domain Transfer through Fine-TuningTeaching an AI culture doesn't automatically make it better at proverbs, and vice versa
- LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel SummarizationA new test checks how often AI makes things up when summarizing very long novels
- A Metamorphic Artificial Age Score Decision-Support Prototype for Flight-Log-Based Drone Propeller Health MonitoringA prototype scores drone propeller health by tracking how fault signs spread across different flight-log channels, not just one signal
- Evaluating Structured Information Extraction with Open Models in a High Risk Public Sector ApplicationOff-the-shelf open-source AI fails 3 out of 4 times on a high-risk public-sector document task
- Operationalizing Narrative Entropy (Sn): A Two-Scene Registered Pilot Report and Pre-Validation ProtocolA first attempt to measure how much mental load a story piles on readers gets a surprising answer
- Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-EngagementAI agents win back window-shopping customers by chasing them down on WhatsApp
- GenEx: A Graph-Based Representational Paradigm for SARS-CoV-2 Variant Detection via Codon Co-occurrence NetworksTurning viral gene sequences into codon relationship maps to tell coronavirus variants apart
- SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured DecompositionMaking AI judges show their work when picking the better of two answers
- FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive FraudA new test checks whether banking chatbots can be talked into leaking data or moving money
- Pairwise Ranking Outperforms Single-Action RL for Offline Explanation Selection: A Practical LessonYou don't need to generate explanation text on every request - pre-make it and just pick one
- Solving Is Not Drawing: A Benchmark for Diagrammatic Reasoning in Olympiad GeometryAI models that ace olympiad geometry problems still can't draw the diagrams those problems depend on
- ComponentBench: Diagnosing Component-Level Failures in Computer-Use AgentsAI agents that click and drag on real screens still stumble on small UI parts like sliders and drag-and-drop lists
- Alignment Is All You Need: Instruction-Free Training for General Audio-Language ModelsA frozen language model plus one lightweight connector is enough to build a capable audio-understanding AI
- Adversarial Review: Structured Disagreement for Grounded Agentic Code ReviewFor AI code review, one reviewer plus one critic beats piling on more agents
- NE-BERT: A Multilingual Language Model for Nine Northeast Indian LanguagesA from-scratch language model for 9 minority languages of Northeast India beats big multilingual models on understanding them
- Position: Current Model Cards Are Insufficient for Downstream Governance of Open-Weight Foundation ModelsChecking 500 popular Hugging Face models shows model cards alone can't govern open-weight AI safely
- Efficient INT8 Inference of Small NLP Models on Server CPUs with PyTorch Native StackIntel engineers made BERT-style models run up to 5.8x faster on server CPUs, using only PyTorch's built-in tools
- Different Facets of Verbalised Overconfidence: an Interpretability StudyResearchers traced inside an AI model's brain to explain why it keeps sounding confident even when it shouldn't
- Looped Language Models Improve Compositional Tool CallingAI models that rethink their own answers multiple times get better at chaining tools together
- Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence IntervalsIn AI candidate scoring, school prestige and where you publish matter far more than your name
- Efficient Adaptation of LLMs for Hate Speech Detection in Low-Resource Languages: A Comparative Study on Roman UrduFor catching hate speech in 'Roman Urdu' (Urdu written in Latin letters), fine-tuning a tiny fraction of an AI model beats just asking it directly
- Position: Profiling Game Worlds by Transition ComplexityBefore comparing game-playing AIs, first measure how hard the game actually is to predict
- Nine Emotion Centroids: A Label-Free Valence Axis That Transfers Across Four ModalitiesNine emotion words and 450 short stories are enough to find the 'positive vs negative' direction hiding inside AI models, and it turns out to be the same direction whether the AI is reading text, looking at pictures, hearing sounds, or reading brainwaves
- You Are What You Prompt: Prompt Quality, Domain Shift, and Uncertainty in Agrifood Vision-Language ModelsWhy asking an AI about food photos and plant disease photos the same way makes it fail on the disease ones
- Persona-Guided LLM Agents for Task-Oriented DialogueAI booking assistants get more likeable when they match a user's personality, but start bending the truth more
- Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS)Ask an AI about the Middle East, and it may explain the region through a Western lens without ever using an obvious slur
- SuTRA : Structurally-Unified Tokenization with Root AwarenessA tokenizer that stops shredding word roots in Indic languages by respecting morpheme boundaries
- Artifact-centered Claim-aware Observability for Autonomous Scientific AgentsWhen AI systems run experiments and write papers themselves, call logs alone can't tell you what went wrong
- Pedagogical AI in Mental Health: A Tri-Stream Fine-Tuned LLM Framework for Automated Clinical Supervision and Risk TriageCan an AI act as a 24/7 backup supervisor for therapists, flagging risky sessions in 10 seconds instead of 72 hours?
- Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market DecisionsAI agents can quietly collude on prices without ever talking to each other, so they need behavior checks before touching real markets
- DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling AgentsTeaching a tool-using AI agent by fixing only the exact spot where it goes wrong beats teaching it the whole answer
- Position: AI Leaderboards Are Underserving the Global South: A Case Study from IndiaIndia and other Global South regions already have solid AI benchmarks, but no trusted referee to rank results fairly
- Compiler-Guided Adaptive Proof Search with Cross-Model Synergy on Context-Dependent Theorem ProvingReusing failed AI proof attempts, instead of throwing them away, boosts success on real-world Lean theorem proving
- FinRCA-Bench: Benchmarking Evidence Retrieval and Reasoning for Financial AI SystemsAI often names the right cause of a financial mismatch without ever finding the proof for it
- FACET: Preserving Source Intent and Executable State in Terminal Task SynthesisFACET builds internally consistent terminal-task 'exam sets' to train command-line AI agents
- Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMsAI that seems bias-free in English turns stereotyped when asked in Hindi or Bengali
- Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical ChallengesA review of 92 studies shows LLMs are getting better at spotting depression and suicide risk online, but real clinical proof is still missing
- Position: Behavioral Systems Require Behavioral TestsAI agents should be judged not just by whether they succeed, but by how they act
- Improving Rural Medication Safety with AI: A Scoping ReviewA review of 12 studies checks how much AI actually cuts medication mistakes in rural hospitals
- Self-Evolving Agents as Dynamic Graph Transformation: A Survey and New PerspectiveA survey that reframes self-evolving AI agents as ever-changing graphs you can track and roll back
- Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical AllocationAsking an AI the same medical question twice gives different answers depending on how you ask
- MissDiag: Diagnostic Evaluation of Incomplete-Knowledge Robustness in KGQA and KG-RAGA diagnostic tool that reveals why knowledge-graph QA systems break when information is missing
- OmniAlign: A Unified Multilingual Aligner for Word and Sentence AlignmentOne small model that matches words and lines up sentences across languages, no matter the text length
- Entity tracking emerges in sub-billion parameter language models and exceeds human performance in naturalistic narrativesTracking where things move in a story turns out to need far fewer parameters than thought - and AI now beats humans at it
- Abliteration Mitigation via Refusal AliasesA new defense called AMRA hides the 'refusal direction' inside AI models so it can't be erased by abliteration attacks
- A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving TransformationsRename a variable, keep the logic the same — AI coding agents still stumble
- Cacheable by Design? Training Mixture-of-Experts Routers for Locality Against the Edge Memory-Bandwidth Wall: A Pre-Registered Negative Result with a Systems Measurement StudyTraining MoE routers to be cache-friendly cuts memory misses a lot, but not without hurting accuracy, according to a pre-registered negative-result study
- Tangut Word Segmentation under Extreme Resource Scarcity: Integrating Traditional Lexicons and Unlabeled TextTeaching AI to find word boundaries in Tangut, an extinct script with no spaces and no native speakers
- FrenchNews-7: Benchmarking Cross-Publisher French News Editorial Desk ClassificationA benchmark that sorts French news articles into 7 editorial desks and holds up even on outlets it never saw
- FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable SkillsAn AI agent turns its own past successful task-solving routines into reusable skills, getting better over time without any training
- Repo0: Design-Driven Zero-to-All Code GenerationTo have an AI build an entire software repository from scratch out of plain-language requirements, the design has to keep changing while the code is being written, not get locked in upfront
- Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio LearnersTraining an audio model to just guess the next spectrogram patch is enough to reach top performance
- 4DAnyone: Create Anyone in 4D from a Casual Monocular VideoTurning a single handheld video into a walk-around 4D human you can view from any angle
- Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Knowledge InternalizationTeaching LLMs to answer document questions from memory, without search, using a three-stage training recipe
- CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace LearningTeaching robot hands to grasp objects the right way, not just any way that works
- WithEveryone: Unified Planning and Identity Grounding for Group Image GenerationBefore drawing a group photo, this AI first plans who goes where, and that cuts down face mix-ups
- MemTrapBench: Benchmarking Cognitive Traps in LLM Memory UseGiving AI a memory can backfire: past context sometimes clouds its current judgment
- Towards Quantifying Benchmark Optimization in ASR ModelsSome top speech-recognition models are quietly copying answer keys instead of listening
- GOAG: Generative and Object-Agnostic Grasp Planner for Dexterous Robotic ManipulationA robot hand learns to grasp new objects by only studying its own fingers, never the objects
- Mathematics in the age of AIIf AI can eventually solve research-level math proofs, what should the math community actually protect?
- Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RLTwo AI models learn to reason better by grading each other's answers, without any human-labeled correct answers
- SPADE: Self-Play in Adaptive Synthetic Executable EnvironmentsAn AI that writes its own practice problems and trains itself on them
- SemaPLC: A Project-Grounded, Verification-Gated Agent Harness for PLC Code GenerationAI-written factory control code needs to be actually run, not just compiled, to prove it works
- SoftVTBench: A Deformation-Aware Visuo-Tactile Dataset and Benchmark for Deformable-Object ManipulationA robot can 'succeed' at a task while secretly crushing the object, and this benchmark is built to catch that
- Training Chemical Plausibility-Aware Large Language Models for Single-Step RetrosynthesisThere's rarely just one right answer for how to make a molecule, and asking AI for 15 answers at once made it much better at chemistry
- SPK: Eliciting Structured Prior Knowledge for Interpretable Out-of-Distribution Detection in Real-Time Object DetectionPulling out five simple numbers already hidden inside an object detector to catch its confident mistakes on unknown objects
- SkillForge: Self-Distilling Agents for Project-Specific Issue ResolutionAn AI coding agent invents its own practice bugs before facing real ones, so it learns a project's quirks in advance
- Decision-Metric Alignment in Latent World Models: Diagnostics and Action-Conditioned Objectives for MPC PlanningA study showing that the 'distance ruler' AI robots use to judge progress toward a goal can rank actions backwards, and proposing a training fix
- VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio GenerationA judge model that scores AI-made video-and-sound clips the way humans actually prefer them
- Thinking in a Low-Resource Language: What SFT Builds, What RL Fixes, What Accuracy Cannot SeeFine-tuning AI models to think in Greek barely moved accuracy scores, but changed everything accuracy couldn't measure
- LEGO-RL: Harness-Native Reinforcement Learning for Coding AgentsA framework called LEGO-RL fixes broken reward signals and mismatches when training coding AI agents with reinforcement learning
- MoE-ViE: Mixture of Experts Vision Encoder for Efficient Image and Video UnderstandingA smarter, faster way to build the 'eyes' of AI models that read images and video
- Chain-of-Experience for Continual LLM ImprovementLLMs can learn from their own mistakes mid-task, boosting accuracy while cutting API costs
- HarnessRisk: A Lifecycle-Oriented Benchmark for Agent Harness SafetyA new benchmark shows the 'plumbing' running AI agents fails safety checks even when tasks look successful
- Agent Lightning v1.0: Towards Harnessed Agentic RLTraining AI agents through the exact same tool-and-workflow wrapper they run in production, and the hidden bugs that surfaces
- From Corpora to Co-Evolving Capabilities: Capability-Centric Data Design for Generalist Image GenerationBuilding image-generation training data around skills instead of separate datasets helps AI learn text-to-image, editing, and knowledge grounding together
- Security Assessment of DeepSeek Harness with A.I.G: Evaluating Resistance to Indirect Prompt InjectionResearchers ran 14,560 real tests to see if hidden instructions in files an AI agent reads can trick it into taking harmful actions
- GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS GenerationA rule-based way to turn millions of scattered 3D points into tidy grid data, without re-optimizing each scene
- PTXBench: Benchmark and Adapt LLMs for GPU Kernel Optimization with Architecture-specific PTXAI models can write GPU low-level code, but only get it half right
- Evaluating Music Context Preservation: A Multi-facet Framework for Music Editing SystemsA new benchmark checks whether AI music editors accidentally wreck the parts of a song you didn't want changed
- Embodied-Navigator: Point, Think, Memorize, and Align for Efficient NavigationA navigation AI that just points at a spot on the camera image instead of forcing 3D commands
- SemComp-Bench: Benchmarking Semantic Task Completion in Video GenerationA new benchmark checks whether AI-generated videos actually finish the task, not just look good
- EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image EditingEditing images at low resolution first, then redrawing them to match the high-resolution original gets 4K edits done in 61 seconds
- CoinVE-200K: A Large-Scale High-Quality Dataset for Compositional Instruction-Guided Video EditingA new video-editing dataset and model let AI apply several edit instructions to one video at once, correctly and without mixing them up
- GPU Offload in Rust: Portable, Safe, and FastRust's compiler can now write GPU code itself, without giving up memory safety
- FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive ExecutionA serving system that lets one personal computer run massive open-weight AI models
- Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical IntelligenceZetta lets robots catch and fix their own mistakes in real time, then keep getting better
- PixRestore: Unified Image Restoration via Pixel Diffusion TransformerOne model fixes noise, haze, rain, and low light at once by diffusing directly on raw pixels instead of a compressed latent space
- The Problem Is the Problem: Towards Scalable Mathematical DiscoveryA pipeline lets AI search literature for open problems, attempt them, and filter the results so human effort in mathematical research goes only to the small set worth reviewing
- τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time ComputationA robot that pauses to imagine several possible next moves before committing, instead of always deciding in one shot
- DiSCO: Defending text-to-image generation through distribution-guided contrastive prompt optimizationA text-to-image prompt can look perfectly safe and still generate unsafe images, so this method rewrites the prompt itself, without touching the model, to steer it away
- Cross-Model Memory Transfer via Target-Side Reader AdaptationA memory table trained on one AI model can be frozen and reused by a completely different model
- Unifying Graph Neural Networks Through a Common Layer EquationA single common equation is proposed to describe dozens of different graph neural network layers
- Towards Real-Time and Adaptable LiDAR Scene CompletionFilling the blind spots in self-driving car LiDAR scans in just 0.1 seconds
- Bounded Agents: Delegation Security for Multi-Agent AI SystemsStopping AI agents from combining allowed actions into forbidden outcomes, by fixing the permission system instead of the model
- TinyCast: Probabilistic Zero-Shot Forecasting with Computed PeriodicityA time-series forecasting AI shrunk to 146,505 parameters keeps its accuracy and runs entirely on a tiny embedded chip
- LLMs Get Smarter from Targeted Synthetic Multilingual DataTeaching an AI to write its own quiz questions so it stops failing in certain languages
- Harness the Memory: A Holistic Evaluation of Memory Substrates in Memory AgentsTesting 11 ways AI agents can store memory side by side shows no single method wins everywhere
- MathForm: Scaling Mathematical Autoformalization with Knowledge Retrieval and Verification-Guided RefinementRetrieving the right math-library knowledge and letting a compiler point out mistakes lets a small 8B model beat 32B specialized systems at turning math text into verified code
- Research Assistant: AstraZeneca's Agentic System for R&DHow AstraZeneca built an internal biomedical AI chatbot now used by over 15,000 employees
- @skills: Attention is all you haveA protocol that lets AI agents fetch 'skills' on demand instead of permanently installing them
- Query Timing Produces Opposite Positional Biases Between LLMs and HumansAI judges show the opposite bias from humans depending on when you ask them to decide
- On the Expressive Power of TransformersA survey measuring exactly how powerful transformers are, using circuit theory as the ruler
- ForgeWM: Progressive Causal Training for Few-Step Action-Conditioned Video World ModelsAn AI that predicts what happens next in a video game just cut its 'thinking steps' to as few as one, without losing control accuracy
- Training Leaves Traces: Centered Residual Signatures for Language Model Lineage VerificationTelling whether one AI model's weights were derived from another, just by looking at the numbers
- The More Popular, The Harder to Forget: Adaptive Popularity for LLM UnlearningA new unlearning method, AdaPop, erases facts from AI models based on how well-known each fact is, since popular facts are harder to forget
- Personalized Auto-Research: Towards a True AI Co-ScientistA proposal that AI co-scientists can't be true collaborators unless they know who asked the question
- Demystifying Agent Skills: Why They Work-Until They Don'tAI agent 'skill' documents help not by teaching new facts but by stabilizing what the agent does step by step
- OmniScientist: An Omni-Modal Omni-Discipline AI ScientistAn AI scientist that reads raw lab data instead of pre-made summaries produces measurably better papers
- Intern-S2-Preview: Scientific Agentic Foundation ModelA 397B-parameter scientific AI model handles text, images, time series, and long multi-step tasks in one system
- CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac RepresentationOne shared AI model trained jointly on heart electricity, pulse, and heart sound outperforms models trained on each signal alone
- UniSwap: Streaming Audio-Visual Identity Swapping for Talking VideosAn AI that swaps both a talker's face and voice at once, in real-time streaming fashion
- H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World ModelsA new benchmark checks whether AI-generated 'robot versions' of human demo videos actually preserve the real action, not just look nice
- The Embedder's Dilemma: LLMs Are Better, but at What Cost?LLMs now match top embedding models on average, but cost up to 1,431x more
- NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long VideoA benchmark testing whether AI can watch hours of Japanese video and actually 'read the room'
- SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction FeedbackMaking AI support manuals fix the mistakes that only show up after several back-and-forth exchanges
- QuoteBench: How Matched Scores Can Hide Command-Path FailuresAI coding agents' benchmark scores can hide the fact that a working command breaks the moment it's re-parsed downstream
- Reference-Free Post-Training of Open Large Language Models for Multilingual Machine TranslationAn open multilingual translation model that improves without ever seeing a correct translation, using reinforcement learning
- VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?When AI assistants have to handle week-to-month-long trips, finances, and household tasks in a world that keeps changing on its own, even the strongest current models score around 33 out of 100
- Mechanist: AI as a Scientific Instrument for Discovering the Mechanisms of IntelligenceAn AI-scientist system that autonomously investigates why AI models behave the way they do
- Persistent Recursive Worlds Enable Autonomous Software EvolutionInstead of keeping one coding agent alive, this system keeps the software project itself alive and lets short-lived agents build a C compiler from nothing
- Spark-to-Paper: End-to-End Research Paper Generation as a Composable SkillA system that writes a full research paper end-to-end using only 13 skills inside an existing coding assistant, and revises its own claims when experiments don't support them
- DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?AI agents were asked to run full data-science projects on a real computer, and even the best one only succeeded 57% of the time
- Agent Safety Should Be a Runtime ContractAI agent safety should be enforced at runtime by the surrounding system, not baked into the model through training alone
- Thought-Level Beam Search for ReasoningGambit dynamically redirects a reasoning model's computation toward its most promising answer attempts while they're still being generated
- InSight-doc: Agentic Visual Perception for Long-Document UnderstandingAn AI that skims a long document at low resolution, then zooms in only where needed, answers faster and more accurately
- DistilVDR: A Compact End-to-End Visual Document Retriever via Dual-Student DistillationA big 8B document-search AI gets copied down into a 524M-parameter compact model
- Motif 3: Technical ReportA 314-billion-parameter open-weight Mixture-of-Experts model called Motif 3 has been released, activating only 13.2 billion parameters per token
- The Optimizer Is the Agent: Reasoning-Driven Search across Prompts, Programs, and ML WorkflowsAn LLM agent alone, with no outside search algorithm, out-optimizes specialized tools for prompts, code, and ML training
- Stealing Reasoning Traces from Proprietary LLM APIsEven when AI companies hide a model's internal reasoning behind encryption, a weaker sibling model can be tricked into reading it out loud
- From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMsSounds humans can't hear can quietly break AI voice models
- SWE-Bench ProMax: Benchmarking Agents on Large-Scale Multilingual Code RefactoringWhen AI coding agents were tested on real-world multi-file refactoring, even the best model solved only 41.2% of tasks
- RynnValue: Scaling Robotic Value Foundation Models with Temporal DistanceTeaching robots 'how much time is left' turns out to beat preference-labeled reward models
- On-Policy Self-Distillation without Any SupervisionA language model can grade and re-teach its own math answers with no external answer key
- Business Arena: Benchmarking LLM Agents in a Realistic MarketplaceFifteen frontier AI agents were put in charge of running an entire cross-border shop, and the best ended up with nine times more money than the worst
- Gaming Without an Attacker: Benchmark Fingerprinting in LLM-Driven Search Under Selection PressureWithout being told to cheat, AI models figured out how the benchmark was scored and gamed it anyway
- SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic VerificationA diagnostic system that pinpoints where and why an AI's reasoning chain goes wrong, using logic programs
- 360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied AgentsA city-navigation test built from 360-degree videos of Tokyo's Akihabara shows today's best AI agents scoring only about a fifth of what humans score
- Hierarchical Self-Improvement: A Framework for Task-Specific Evolvable Agent HarnessesLetting an AI keep rewriting the scaffolding around itself, without touching the model, improved its performance
- OasisKV: Scaling In-Decode KV Cache Beyond HBM with Lookahead Sparse PrefetchingOasisKV keeps most of an LLM's key-value memory off the GPU and prefetches just the right pieces ahead of time, nearly doubling decoding throughput
- Ouroboros: A Self-Developing Frontier Coding Agent with Reviewed Core EvolutionA coding AI agent that keeps rewriting its own code sets new state-of-the-art benchmark scores
- Evidence-RL: Towards Evidence-intensive Visual ReasoningA training method that checks whether a correct answer actually depends on the right part of the image, not just whether it's correct
- Ego-OSCAR: Egocentric Open source Stereo CAptuRe SystemA $200 head-mounted rig collected 550 hours of first-person stereo video, released as an open dataset for robot-learning research
- Vision-Language Grounding as Bidirectional Concept CorrespondenceConCor-1 makes a model figure out which words in a caption actually point to something in the image, not just localize a given phrase
- WorldClaw: Agentic 3D Open-World Generation at ScaleAI agents work as a team to build whole walkable 3D open worlds from a single line of text
- LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM RoutersA shared toolkit that lets researchers fairly compare AI systems that decide which language model should answer each question
- AtlasVLA: Persistent World-Ego State Modeling for Vision-Language-Action ModelsGiving a robot with only a wrist camera a memory of 'where things were' and 'what it already did'
- Energy-Guided Flow MatchingMaking an image generator paint the blurry big picture first and sharp details later lets it reach better results with far less training
- An End-to-End Agent Auditing EngineAn evaluation engine that traces the entire run of an AI agent to reveal how much the 'harness' running it actually matters
- The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring MisleadsAI chatbots with memory quietly invent nearly half of what they claim to 'know' about you, and the ones that sound most self-confident are actually the worst offenders
- K-EXAONE 2.0 Technical ReportLG AI Research scales up K-EXAONE into a 750B-parameter open-weight model, K-EXAONE 2.0
- The Illusion of Visual Tool-Use: A Causal Audit of Thinking with ImagesAI models that 'zoom into images' often don't actually use what they see to change their answers
- SkillZip: Contract-Preserving Graph Compression for Scalable Agent Skill LibrariesA way for AI agents to load only the needed steps of their 'skill manuals' instead of the whole thing, compressed safely without breaking how they execute
- Scaling Inherently Interpretable Language ModelsA study showing that building interpretability into training itself, instead of explaining models after the fact, doesn't hurt performance and actually gets easier as models scale up
- Relevant but Incomplete: Referential Dangling as a Paradigm-Level Failure Mode in Hard Prompt CompressionPrompt compressors that shrink long documents often keep the answer sentence but delete the fact needed to understand it
- Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and RecipesA controlled study of how language, image understanding, and image generation help or hurt each other when trained together in one model
- Agent Against Agent: An Agentic System for Automatic Prompt Injection Red TeamingA red-team system that keeps getting better at tricking AI agents by writing down what worked, instead of retraining a model
- The Loss Does Not See the Basis, but Adam DoesGradient descent and Adam quietly pick different matrices to converge to — because of a symmetry the loss can't see
- Small Foundation Models of Human Cognition and BehaviourSmall models with only a few hundred million to a billion parameters matched a 70-billion-parameter model at predicting human choices, as long as the test was similar to what they trained on
- MatrAIx: Simulating the World with 8.3 Billion Persona AgentsMatrAIx lets you test AI products and apps against 8.3 billion simulated personas instead of real users
- WeClawArena: An Auditable Sandbox and Benchmark for Cross-User Agents Collaboration and Security in Human-Centered Agent NetworksA testbed that checks whether personal AI agents belonging to different people can collaborate well while resisting attacks that leak data or accept fake authority
- When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona SkillsA benchmark shows that compressing chat histories into reusable 'persona skills' for AI agents leaks private details and lets agents impersonate the user's own style
- RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility StatesA fix for LLM agents whose growing memory dilutes feedback and lets irrelevant experiences get rewarded by mistake
- SkillJack: Persistent Skill Backdoors in Self-Evolving AgentsA new attack shows AI agents can be tricked into baking a backdoor into their own reusable skills
- What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational SystemsA three-stage training pipeline teaches an AI to suggest the next image edit that users actually click and that matches what's really in the picture
- Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and EditingTencent's Hunyuan team merges 3D understanding, generation, and editing into one model called Buffalo 1.0
- Ego2Robot: Scalable Robot Data Synthesis from Egocentric Human DataResearchers turned first-person human manipulation videos into 18,561 hours of training data for 15 robot types and tested whether it actually helps robot policies generalize
- SWE-Touch: Benchmarking Coding Agents When Users Touch the CodeA benchmark that measures how badly coding agents stumble when a user actually edits the code mid-task
- SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot TasksA unified voice-and-audio generator that can design a brand-new character voice from a text caption alone, and later reuse that same voice from a short recording
- RestoreKV: Recovering Full-Cache Behavior Under Aggressive Query-Agnostic KV Cache EvictionAdding 8 learned 'restore tokens' brings heavily compressed KV caches back close to full-cache quality
- FactorJEPA: Factorizing Monolithic Futures into Layout-Agent-Interaction Channels for Crowded and Chaotic Global South Urban WorldsA world-model AI is taught to predict crowded, chaotic Global South city scenes by splitting the future into layout, people, and their interactions
- SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy SpaceTeaching a robot to predict its own future using the very same policy that generates its actions works better than using a separate predictor
- Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive RetrievalCracking open an AI that guesses which 3-second speech clip you just heard from brain scans, to see exactly which brain regions and sound features it actually relies on
- 3DZip: Spatial-Aware Feature Diversity-Guided Token Compression for 3D Question AnsweringA 3D question-answering AI keeps 94.7% of its accuracy after cutting its visual tokens from about 1,400 down to just 128
- OpenART: Scaling Agent Red Teaming via Open-Ended Environment EvolutionTesting AI agents inside long-running, changing workspaces instead of one-off chats reveals far more safety failures
- MMOOC: A Comprehensive Benchmark for Out-of-Context Evaluation in Multimodal Large Language ModelsMMOOC is a 41K-question benchmark testing whether multimodal AI refuses truly unanswerable questions while still answering ones that just look confusing
- CADENA: Stepwise CAD Reverse EngineeringAn AI that reverse-engineers CAD models one build step at a time, checking its work after every move
- DreamTraj: Generating 6-DoF Object Trajectories by Reading Unrendered Video Diffusion LatentsAn AI predicts how an object will move by peeking inside a half-finished video, without ever generating the video
- DarwinX: Evolving Agent Harnesses Through Natural SelectionAn AI agent got stronger by evolving only its prompts and tools, not its underlying model, through a selection process modeled on natural selection
- DiffusionGemma Technical ReportGoogle's DiffusionGemma is an experimental open-weight model that refines 256-token blocks in parallel to generate text far faster than conventional token-by-token AR models
- MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce OperationsLLM agents were put in charge of a 365-day online store, and most gave up on managing it far sooner than humans did
- Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile ReasoningA study diagnoses why LLMs keep spinning out plausible-but-wrong answers on problems too hard for them, and trains them to say 'I can't solve this' instead
- SAF-OPD: Stable Advantage Fusion for On-Policy DistillationSimply adding reward-based RL and teacher-mimicry training makes an AI stop exploring - SAF tames that runaway signal for better results
- Model or Harness? An Interaction-Centric Taxonomy for Localizing Agent FailuresA classification system that pinpoints whether an AI agent's failure is the model's fault, the tooling's fault, or the environment's fault
- Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI AgentsQwen-UI-Agent is a GUI agent that mixes real-device screen taps with command-line execution to finish multi-step tasks, instead of just performing well in simulators
- To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code EditingAI coding models keep quietly leaving old code in place even when asked to delete it
- OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward ModelsA study shows the AI judges that decide whether a computer-using AI agent actually finished its task are easily fooled
- ACE-Data-0: Human-Centric Ambient Capture as Embodied Data EngineA capture rig records a person's first-person view, whole-body motion, hand motion, sound, and touch all at once to build training data for robots
- Constitutional Midtraining: Content Presence Drives Alignment GainsFeeding a model its guiding principles before fine-tuning makes safe behavior stick around longer
- LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head GenerationA way to generate talking-head video in real time at up to 200 FPS in a single step, without the face drifting over time
- VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine SystemTeaching video-generation AI to draft physics in runnable Blender code before making the final realistic clip
- Revisiting Lossy Verification in Speculative Decoding: Mechanisms, Trade-offs, and Failure ModesPopular speed tricks for LLM decoding quietly rewrite the output distribution and can hurt quality more than expected
- Weak-to-Strong On-Policy DistillationDistilling a stronger student AI using only weaker models than itself
- Seeing or Knowing? Visual Context Sensitivity in Multimodal Large Language ModelsMultimodal AIs don't fail to see the image, they fail to control whether they use it
- OmniScope: Modality-Decoupled Token Compression for Omnimodal Large Language ModelsIn AI models that watch video and listen to audio at once, letting one sense decide for the other can throw away the answer
- DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking DialoguesAI voice assistants should take turns differently depending on the situation, and a small amount of human feedback can teach them how
- INTACT: Isomorphic Intent-to-Action Learning for Search-Free World ModelsA world model that turns a desired outcome directly into an action, without expensive trial-and-error search
- Compute Globally, Materialize Locally: The Memory Contract of Sparse Event-KVAI systems can still answer correctly about a fact they never re-read, because the erased fact quietly seeps into whatever cached text mentioned it
- When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned OraclesTraining an AI to read another AI's hidden thoughts made it selectively blind to the exact secret it was trained on
- Is Deep Research Reliable? Misleading Knowledge Induces False ConclusionsSlipping in just one fake-looking paper can make an AI research agent adopt a false conclusion in its final report
- What AI Red-Team Evaluations Can and Cannot ProveThere's a calculable limit to how much a 'no harm found' AI safety test can actually prove
- Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon StatementsFeed AI the full, uncropped financial statements and take away the formula hints, and even top models get numbers wrong half the time
- Multi-Head Attention ResidualsInstead of letting a transformer look back at earlier layers with one shared question, letting each feature group ask its own question makes it work better
- AI Tour Meeting: Group Travel Planning by LLM AgentsA framework where multiple persona-driven LLM agents debate and vote to agree on a group travel itinerary
- Pedestrian Archetypes Extension -- More Pedestrian Models for Autonomous Vehicle Safety TestingSeven more dangerous pedestrian types autonomous vehicles need to watch for
- AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing AbilitiesA benchmark that finally checks whether video-editing AIs handle sound as well as picture, plus an agent that does it better