03 / ARTICLE LEDGER

关联文章

按发布时间倒序

  1. Agentic Uncertainty Reveals Agentic Overconfidence 2026-02-09 · ARXIV
  2. Shared LoRA Subspaces for almost Strict Continual Learning 2026-02-06 · ARXIV
  3. Pseudo-Invertible Neural Networks 2026-02-06 · ARXIV
  4. PhysicsAgentABM: Physics-Guided Generative Agent-Based Modeling 2026-02-06 · ARXIV
  5. Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory 2026-02-06 · ARXIV
  6. DyTopo: Dynamic Topology Routing for Multi-Agent Reasoning via Semantic Matching 2026-02-06 · ARXIV
  7. DFlash: Block Diffusion for Flash Speculative Decoding 2026-02-06 · ARXIV
  8. Curiosity is Knowledge: Self-Consistent Learning and No-Regret Optimization with Active Inference 2026-02-06 · ARXIV
  9. CommCP: Efficient Multi-Agent Coordination via LLM-Based Communication with Conformal Prediction 2026-02-06 · ARXIV
  10. Can vision language models learn intuitive physics from interaction? 2026-02-06 · ARXIV
  11. AP-OOD: Attention Pooling for Out-of-Distribution Detection 2026-02-06 · ARXIV
  12. Wedge Sampling: Efficient Tensor Completion with Nearly-Linear Sample Complexity 2026-02-06 · ARXIV
  13. RRAttention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context Inference 2026-02-06 · ARXIV
  14. Exact Recovery in the Data Block Model 2026-02-06 · ARXIV
  15. DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders 2026-02-06 · ARXIV
  16. Constrained Group Relative Policy Optimization 2026-02-06 · ARXIV
  17. Subliminal Effects in Your Data: A General Mechanism via Log-Linearity 2026-02-05 · ARXIV
  18. Rethinking the Trust Region in LLM Reinforcement Learning 2026-02-05 · ARXIV
  19. Reinforced Attention Learning 2026-02-05 · ARXIV
  20. Protein Autoregressive Modeling via Multiscale Structure Generation 2026-02-05 · ARXIV
  21. Multi-layer Cross-Attention is Provably Optimal for Multi-modal In-context Learning 2026-02-05 · ARXIV
  22. Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism 2026-02-05 · ARXIV
  23. CRoSS: A Continual Robotic Simulation Suite for Scalable Reinforcement Learning with High Task Diversity and Realistic Physics Simulation 2026-02-05 · ARXIV
  24. CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation 2026-02-05 · ARXIV
  25. Contrastive Continual Learning for Model Adaptability in Internet of Things 2026-02-05 · ARXIV
  26. Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL 2026-02-04 · ARXIV
  27. Robust Intervention Learning from Emergency Stop Interventions 2026-02-04 · ARXIV
  28. PrevizWhiz: Combining Rough 3D Scenes and 2D Video to Guide Generative Video Previsualization 2026-02-04 · ARXIV
  29. PLATE: Plasticity-Tunable Efficient Adapters for Geometry-Aware Continual Learning 2026-02-04 · ARXIV
  30. Parallel-Probe: Towards Efficient Parallel Thinking via 2D Probing 2026-02-04 · ARXIV
  31. AutoFigure: Generating and Refining Publication-Ready Scientific Illustrations 2026-02-04 · ARXIV
  32. Accelerating Scientific Research with Gemini: Case Studies and Common Techniques 2026-02-04 · ARXIV
  33. Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability 2026-02-03 · ARXIV
  34. RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System 2026-02-03 · ARXIV
  35. Reward-free Alignment for Conflicting Objectives 2026-02-03 · ARXIV
  36. RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents 2026-02-03 · ARXIV
  37. PixelGen: Pixel Diffusion Beats Latent Diffusion with Perceptual Loss 2026-02-03 · ARXIV
  38. MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents 2026-02-03 · ARXIV
  39. MEG-XL: Data-Efficient Brain-to-Text via Long-Context Pre-Training 2026-02-03 · ARXIV
  40. Flow Policy Gradients for Robot Control 2026-02-03 · ARXIV
  41. Expanding the Capabilities of Reinforcement Learning via Text Feedback 2026-02-03 · ARXIV
  42. AgentRx: Diagnosing AI Agent Failures from Execution Trajectories 2026-02-03 · ARXIV
  43. Sparse Reward Subsystem in Large Language Models 2026-02-03 · ARXIV
  44. Scalable Random Wavelet Features: Efficient Non-Stationary Kernel Approximation with Convergence Guarantees 2026-02-03 · ARXIV
  45. Reliable Use of Lemmas via Eligibility Reasoning and Section$-$Aware Reinforcement Learning 2026-02-03 · ARXIV
  46. Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning 2026-02-03 · ARXIV
  47. Optimal Decision-Making Based on Prediction Sets 2026-02-03 · ARXIV
  48. How RLHF Amplifies Sycophancy 2026-02-03 · ARXIV
  49. HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving 2026-02-03 · ARXIV
  50. Error Taxonomy-Guided Prompt Optimization 2026-02-03 · ARXIV