03 / ARTICLE LEDGER

关联文章

按发布时间倒序

  1. Mechanistic Reaction Prediction via Discrete Flow Matching on Graph-Structured Electron Occupation 2026-08-30 · ARXIV
  2. RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution 2026-08-29 · ARXIV
  3. From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench 2026-08-29 · ARXIV
  4. SWE-Prime: Fewer Trajectories, Better Performance 2026-08-29 · ARXIV
  5. WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution 2026-08-29 · ARXIV
  6. CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes 2026-08-28 · ARXIV
  7. A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training 2026-08-28 · ARXIV
  8. VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning 2026-08-27 · ARXIV
  9. Automatic Model Card Generation Using an LLM 2026-08-27 · ARXIV
  10. Structurally-bounded Agentic Graph Exploration for Evidence-Grounded Scholarly DeepSearch 2026-08-27 · ARXIV
  11. Strictly Causal Streaming Video Anomaly Detection with a Theoretically-Grounded State-Space Core 2026-08-26 · ARXIV
  12. MDTE: Minority-Aware Diffusion over Temporal Edge Events for Imbalanced Node Classification 2026-08-26 · ARXIV
  13. Constrained Entity Selection under Partial Knowledge for LLM-Based Knowledge Graph QA 2026-08-26 · ARXIV
  14. A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments 2026-08-26 · ARXIV
  15. Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows 2026-08-26 · ARXIV
  16. LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training 2026-08-26 · ARXIV
  17. FedV-KGQA: Multi-Hop Question Answering over Vertically Partitioned Knowledge Graphs 2026-08-26 · ARXIV
  18. BrowserForge: Scaling Web Episode via Parallel Browser Sandboxes 2026-08-26 · ARXIV
  19. Learning Whom to Trust : Decision-Generated Credibility in Social Learning 2026-08-26 · ARXIV
  20. Parameterized Complexity of $L_p$-Lipschitz Constants for Input Convex Neural Networks and $L_p$-Norm Maximization over Zonotopes 2026-08-26 · ARXIV
  21. SPO++: Stream-Aligned Policy Optimization for Asynchronous Agentic RL 2026-08-26 · ARXIV
  22. Recursive Experiential-Working Memory Evolution for Long-Horizon Agent Harnesses 2026-08-26 · ARXIV
  23. Mitigating Reasoning-Induced Misalignment via Safety-Direction Penalty 2026-08-26 · ARXIV
  24. When Names Cross Scripts: A Source-Grounded Benchmark for Historical Entity Reconciliation in the Mongol World 2026-08-26 · ARXIV
  25. The Measurement Revolution? Credible Measurement and Inference in the Age of AI 2026-08-26 · ARXIV
  26. EarthVerse: Benchmarking Scientific Agents Across Dynamic Earth Systems and Natural Hazards 2026-08-26 · ARXIV
  27. Correcting a learned physical invariant improves world-model rollouts 2026-08-26 · ARXIV
  28. Predicting Multiple Clinical Outcomes Related to Functional Recovery and Social Isolation Among Older Adults After Lower-Limb Fracture or Hip Replacement 2026-08-26 · ARXIV
  29. Interpretable AI with Local Distillation 2026-08-25 · ARXIV
  30. The Interaction Tax: When Communication Erases Diversity in Multi-Agent Teams 2026-08-25 · ARXIV
  31. How AI Assistance Affects Human Skill Development: A Study of Learning with Logic Puzzles 2026-08-25 · ARXIV
  32. ConvergeFlow: Language Flow with Provable Convergence to Token Embeddings 2026-08-25 · ARXIV
  33. Prime Agent: A Self-Improving RLM Harness 2026-08-25 · ARXIV
  34. Provably adaptive sampling with uniform and remasking discrete diffusion models 2026-08-25 · ARXIV
  35. Physics-Constrained Deep Learning Model for Contactless Blood Pressure Monitoring from Triaxial Bodyseismography 2026-08-25 · ARXIV
  36. EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings 2026-08-25 · ARXIV
  37. SWE Refactor Bench: Can Coding Agents Complete a Long-Horizon, Whole-Repository Stack Migration? 2026-08-25 · ARXIV
  38. ReWorld: An Interactive World Model with Long-Horizon Memory 2026-08-25 · ARXIV
  39. How to Train a Critic Stably and Efficiently 2026-08-25 · ARXIV
  40. OmniScientist: An Omni-Modal Omni-Discipline AI Scientist 2026-08-14 · ARXIV
  41. AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design 2026-08-14 · ARXIV
  42. One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL 2026-08-14 · ARXIV
  43. Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams 2026-08-14 · ARXIV
  44. A Cascaded Unsupervised-Supervised NLP Pipeline for Detecting Accusatory Language in Public Procurement 2026-08-14 · ARXIV
  45. Earth observation embeddings are effective sub-grid descriptors for probabilistic weather downscaling 2026-08-14 · ARXIV
  46. Convergent Detour Hijacking: Task-Preserving Resource Amplification in Skill-Based LLM Agents 2026-08-14 · ARXIV
  47. A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery 2026-08-14 · ARXIV
  48. Structural Silence: When AI Infrastructure Fails Speakers of Underrepresented Languages 2026-08-13 · ARXIV
  49. VAKRA: Evaluating Multi-Hop Reasoning Across APIs and Retrieval Under Tool-Use Policies 2026-08-13 · ARXIV
  50. Large Language Model-Driven Small-Capitalization Trading: Integrating Financial News Sentiment, Macroeconomic Indicators, and Technical Signals 2026-08-13 · ARXIV