03 / ARTICLE LEDGER
关联文章
按发布时间倒序
- Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation 2026-08-30 · ARXIV
- RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution 2026-08-29 · ARXIV
- From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench 2026-08-29 · ARXIV
- SWE-Prime: Fewer Trajectories, Better Performance 2026-08-29 · ARXIV
- WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution 2026-08-29 · ARXIV
- CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes 2026-08-28 · ARXIV
- A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training 2026-08-28 · ARXIV
- VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning 2026-08-27 · ARXIV
- Automatic Model Card Generation Using an LLM 2026-08-27 · ARXIV
- Structurally-bounded Agentic Graph Exploration for Evidence-Grounded Scholarly DeepSearch 2026-08-27 · ARXIV
- Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Space Core 2026-08-26 · ARXIV
- MDTE: Minority-Aware Diffusion over Temporal Edge Events for Imbalanced Node Classification 2026-08-26 · ARXIV
- Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA 2026-08-26 · ARXIV
- A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments 2026-08-26 · ARXIV
- Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows 2026-08-26 · ARXIV
- LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training 2026-08-26 · ARXIV
- FedV-KGQA: Multi-Hop Question Answering over Vertically Partitioned Knowledge Graphs 2026-08-26 · ARXIV
- BrowserForge: Scaling Web Episode via Parallel Browser Sandboxes 2026-08-26 · ARXIV
- Learning Whom to Trust : Decision-Generated Credibility in Social Learning 2026-08-26 · ARXIV
- Parameterized Complexity of $L_p$-Lipschitz Constants for Input Convex Neural Networks and $L_p$-Norm Maximization over Zonotopes 2026-08-26 · ARXIV
- SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL 2026-08-26 · ARXIV
- Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses 2026-08-26 · ARXIV
- Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty 2026-08-26 · ARXIV
- When Names Cross Scripts: A Source-Grounded Benchmark for Historical Entity Reconciliation in the Mongol World 2026-08-26 · ARXIV
- The Measurement Revolution? Credible Measurement and Inference in the Age of AI 2026-08-26 · ARXIV
- EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards 2026-08-26 · ARXIV
- Correcting a learned physical invariant improves world-model rollouts 2026-08-26 · ARXIV
- Predicting Multiple Clinical Outcomes Related to Functional Recovery and Social Isolation Among Older Adults After Lower-Limb Fracture or Hip Replacement 2026-08-26 · ARXIV
- Interpretable AI with Local Distillation 2026-08-25 · ARXIV
- The Interaction Tax: When Communication Erases Diversity in Multi-Agent Teams 2026-08-25 · ARXIV
- How AI Assistance Affects Human Skill Development: A Study of Learning with Logic Puzzles 2026-08-25 · ARXIV
- ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings 2026-08-25 · ARXIV
- Prime Agent: A Self-Improving RLM Harness 2026-08-25 · ARXIV
- Provably adaptive sampling with uniform and remasking discrete diffusion models 2026-08-25 · ARXIV
- Physics-Constrained Deep Learning Model for Contactless Blood Pressure Monitoring from Triaxial Bodyseismography 2026-08-25 · ARXIV
- EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings 2026-08-25 · ARXIV
- SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? 2026-08-25 · ARXIV
- ReWorld: An Interactive World Model with Long-Horizon Memory 2026-08-25 · ARXIV
- How to Train a Critic Stably and Efficiently 2026-08-25 · ARXIV
- OmniScientist: An Omni-Modal Omni-Discipline AI Scientist 2026-08-14 · ARXIV
- AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design 2026-08-14 · ARXIV
- One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL 2026-08-14 · ARXIV
- Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams 2026-08-14 · ARXIV
- A Cascaded Unsupervised-Supervised NLP Pipeline for Detecting Accusatory Language in Public Procurement 2026-08-14 · ARXIV
- Earth observation embeddings are effective sub-grid descriptors for probabilistic weather downscaling 2026-08-14 · ARXIV
- Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents 2026-08-14 · ARXIV
- A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery 2026-08-14 · ARXIV
- Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages 2026-08-13 · ARXIV
- VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies 2026-08-13 · ARXIV
- Large Language Model-Driven Small-Capitalization Trading: Integrating Financial News Sentiment, Macroeconomic Indicators, and Technical Signals 2026-08-13 · ARXIV