03 / ARTICLE LEDGER
关联文章
按发布时间倒序
- Knowledge-Embedded Latent Projection for Robust Representation Learning 2026-02-19 · ARXIV
- Causality is Key for Interpretability Claims to Generalise 2026-02-19 · ARXIV
- Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents 2026-02-19 · ARXIV
- Are Object-Centric Representations Better At Compositional Generalization? 2026-02-19 · ARXIV
- Task-Agnostic Continual Learning for Chest Radiograph Classification 2026-02-18 · ARXIV
- Stabilizing Test-Time Adaptation of High-Dimensional Simulation Surrogates via D-Optimal Statistics 2026-02-18 · ARXIV
- Solving Parameter-Robust Avoid Problems with Unknown Feasibility using Reinforcement Learning 2026-02-18 · ARXIV
- Operationalising the Superficial Alignment Hypothesis via Task Complexity 2026-02-18 · ARXIV
- Dex4D: Task-Agnostic Point Track Policy for Sim-to-Real Dexterous Manipulation 2026-02-18 · ARXIV
- Developing AI Agents with Simulated Data: Why, what, and how? 2026-02-18 · ARXIV
- CrispEdit: Low-Curvature Projections for Scalable Non-Destructive LLM Editing 2026-02-18 · ARXIV
- Avey-B 2026-02-18 · ARXIV
- Text Style Transfer with Parameter-efficient LLM Finetuning and Round-trip Translation 2026-02-17 · ARXIV
- Symmetry in language statistics shapes the geometry of model representations 2026-02-17 · ARXIV
- Scaling Beyond Masked Diffusion Language Models 2026-02-17 · ARXIV
- Rethinking Diffusion Models with Symmetries through Canonicalization with Applications to Molecular Graph Generation 2026-02-17 · ARXIV
- Long Context, Less Focus: A Scaling Gap in LLMs Revealed through Privacy and Personalization 2026-02-17 · ARXIV
- Hunt Globally: Deep Research AI Agents for Drug Asset Scouting in Investing, Business Development, and Search & Evaluation 2026-02-17 · ARXIV
- Efficient Sampling with Discrete Diffusion Models: Sharp and Adaptive Guarantees 2026-02-17 · ARXIV
- Cold-Start Personalization via Training-Free Priors from Structured World Models 2026-02-17 · ARXIV
- BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames 2026-02-17 · ARXIV
- When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift 2026-02-17 · ARXIV
- UniWeTok: An Unified Binary Tokenizer with Codebook Size $\mathit{2^{128}}$ for Unified Multimodal Large Language Model 2026-02-17 · ARXIV
- Process-Supervised Multi-Agent Reinforcement Learning for Reliable Clinical Reasoning 2026-02-17 · ARXIV
- Knowing When Not to Answer: Abstention-Aware Scientific Reasoning 2026-02-17 · ARXIV
- Index Light, Reason Deep: Deferred Visual Ingestion for Visual-Dense Document Question Answering 2026-02-17 · ARXIV
- GPT-5 vs Other LLMs in Long Short-Context Performance 2026-02-17 · ARXIV
- Deep Dense Exploration for LLM Reinforcement Learning via Pivot-Driven Resampling 2026-02-17 · ARXIV
- Semantic Chunking and the Entropy of Natural Language 2026-02-16 · ARXIV
- Realistic Face Reconstruction from Facial Embeddings via Diffusion Models 2026-02-16 · ARXIV
- In-Context Autonomous Network Incident Response: An End-to-End Large Language Model Agent Approach 2026-02-16 · ARXIV
- Improved Regret Guarantees for Online Mirror Descent using a Portfolio of Mirror Maps 2026-02-16 · ARXIV
- Imitating What Works: Simulation-Filtered Modular Policy Learning from Human Videos 2026-02-16 · ARXIV
- CoPE-VideoLM: Codec Primitives For Efficient Video Language Models 2026-02-16 · ARXIV
- Asynchronous Verified Semantic Caching for Tiered LLM Architectures 2026-02-16 · ARXIV
- Creative Ownership in the Age of AI 2026-02-14 · ARXIV
- UniT: Unified Multimodal Chain-of-Thought Test-time Scaling 2026-02-13 · ARXIV
- Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action Alignment 2026-02-13 · ARXIV
- On-Policy Context Distillation for Language Models 2026-02-13 · ARXIV
- MonarchRT: Efficient Attention for Real-Time Video Generation 2026-02-13 · ARXIV
- CM2: Reinforcement Learning with Checklist Rewards for Multi-Turn and Multi-Step Agentic Tool Use 2026-02-13 · ARXIV
- AttentionRetriever: Attention Layers are Secretly Long Document Retrievers 2026-02-13 · ARXIV
- Agentic Test-Time Scaling for WebAgents 2026-02-13 · ARXIV
- The Pensieve Paradigm: Stateful Language Models Mastering Their Own Context 2026-02-13 · ARXIV
- Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated Penalty 2026-02-13 · ARXIV
- P-GenRM: Personalized Generative Reward Model with Test-time User-based Scaling 2026-02-13 · ARXIV
- On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial Coverage 2026-02-13 · ARXIV
- Meta-Sel: Efficient Demonstration Selection for In-Context Learning via Supervised Meta-Learning 2026-02-13 · ARXIV
- Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation 2026-02-13 · ARXIV
- KAN-FIF: Spline-Parameterized Lightweight Physics-based Tropical Cyclone Estimation on Meteorological Satellite 2026-02-13 · ARXIV