03 / ARTICLE LEDGER
关联文章
按发布时间倒序
- Agentic Uncertainty Reveals Agentic Overconfidence 2026-02-09 · ARXIV
- Shared LoRA Subspaces for almost Strict Continual Learning 2026-02-06 · ARXIV
- Pseudo-Invertible Neural Networks 2026-02-06 · ARXIV
- PhysicsAgentABM: Physics-Guided Generative Agent-Based Modeling 2026-02-06 · ARXIV
- Learning Query-Aware Budget-Tier Routing for Runtime Agent Memory 2026-02-06 · ARXIV
- DyTopo: Dynamic Topology Routing for Multi-Agent Reasoning via Semantic Matching 2026-02-06 · ARXIV
- DFlash: Block Diffusion for Flash Speculative Decoding 2026-02-06 · ARXIV
- Curiosity is Knowledge: Self-Consistent Learning and No-Regret Optimization with Active Inference 2026-02-06 · ARXIV
- CommCP: Efficient Multi-Agent Coordination via LLM-Based Communication with Conformal Prediction 2026-02-06 · ARXIV
- Can vision language models learn intuitive physics from interaction? 2026-02-06 · ARXIV
- AP-OOD: Attention Pooling for Out-of-Distribution Detection 2026-02-06 · ARXIV
- Wedge Sampling: Efficient Tensor Completion with Nearly-Linear Sample Complexity 2026-02-06 · ARXIV
- RRAttention: Dynamic Block Sparse Attention via Per-Head Round-Robin Shifts for Long-Context Inference 2026-02-06 · ARXIV
- Exact Recovery in the Data Block Model 2026-02-06 · ARXIV
- DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse Autoencoders 2026-02-06 · ARXIV
- Constrained Group Relative Policy Optimization 2026-02-06 · ARXIV
- Subliminal Effects in Your Data: A General Mechanism via Log-Linearity 2026-02-05 · ARXIV
- Rethinking the Trust Region in LLM Reinforcement Learning 2026-02-05 · ARXIV
- Reinforced Attention Learning 2026-02-05 · ARXIV
- Protein Autoregressive Modeling via Multiscale Structure Generation 2026-02-05 · ARXIV
- Multi-layer Cross-Attention is Provably Optimal for Multi-modal In-context Learning 2026-02-05 · ARXIV
- Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism 2026-02-05 · ARXIV
- CRoSS: A Continual Robotic Simulation Suite for Scalable Reinforcement Learning with High Task Diversity and Realistic Physics Simulation 2026-02-05 · ARXIV
- CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation 2026-02-05 · ARXIV
- Contrastive Continual Learning for Model Adaptability in Internet of Things 2026-02-05 · ARXIV
- Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL 2026-02-04 · ARXIV
- Robust Intervention Learning from Emergency Stop Interventions 2026-02-04 · ARXIV
- PrevizWhiz: Combining Rough 3D Scenes and 2D Video to Guide Generative Video Previsualization 2026-02-04 · ARXIV
- PLATE: Plasticity-Tunable Efficient Adapters for Geometry-Aware Continual Learning 2026-02-04 · ARXIV
- Parallel-Probe: Towards Efficient Parallel Thinking via 2D Probing 2026-02-04 · ARXIV
- AutoFigure: Generating and Refining Publication-Ready Scientific Illustrations 2026-02-04 · ARXIV
- Accelerating Scientific Research with Gemini: Case Studies and Common Techniques 2026-02-04 · ARXIV
- Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability 2026-02-03 · ARXIV
- RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System 2026-02-03 · ARXIV
- Reward-free Alignment for Conflicting Objectives 2026-02-03 · ARXIV
- RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents 2026-02-03 · ARXIV
- PixelGen: Pixel Diffusion Beats Latent Diffusion with Perceptual Loss 2026-02-03 · ARXIV
- MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents 2026-02-03 · ARXIV
- MEG-XL: Data-Efficient Brain-to-Text via Long-Context Pre-Training 2026-02-03 · ARXIV
- Flow Policy Gradients for Robot Control 2026-02-03 · ARXIV
- Expanding the Capabilities of Reinforcement Learning via Text Feedback 2026-02-03 · ARXIV
- AgentRx: Diagnosing AI Agent Failures from Execution Trajectories 2026-02-03 · ARXIV
- Sparse Reward Subsystem in Large Language Models 2026-02-03 · ARXIV
- Scalable Random Wavelet Features: Efficient Non-Stationary Kernel Approximation with Convergence Guarantees 2026-02-03 · ARXIV
- Reliable Use of Lemmas via Eligibility Reasoning and Section$-$Aware Reinforcement Learning 2026-02-03 · ARXIV
- Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning 2026-02-03 · ARXIV
- Optimal Decision-Making Based on Prediction Sets 2026-02-03 · ARXIV
- How RLHF Amplifies Sycophancy 2026-02-03 · ARXIV
- HERMES: A Holistic End-to-End Risk-Aware Multimodal Embodied System with Vision-Language Models for Long-Tail Autonomous Driving 2026-02-03 · ARXIV
- Error Taxonomy-Guided Prompt Optimization 2026-02-03 · ARXIV