03 / ARTICLE LEDGER

关联文章

按发布时间倒序

  1. Knowledge-Embedded Latent Projection for Robust Representation Learning 2026-02-19 · ARXIV
  2. Causality is Key for Interpretability Claims to Generalise 2026-02-19 · ARXIV
  3. Calibrate-Then-Act: Cost-Aware Exploration in LLM Agents 2026-02-19 · ARXIV
  4. Are Object-Centric Representations Better At Compositional Generalization? 2026-02-19 · ARXIV
  5. Task-Agnostic Continual Learning for Chest Radiograph Classification 2026-02-18 · ARXIV
  6. Stabilizing Test-Time Adaptation of High-Dimensional Simulation Surrogates via D-Optimal Statistics 2026-02-18 · ARXIV
  7. Solving Parameter-Robust Avoid Problems with Unknown Feasibility using Reinforcement Learning 2026-02-18 · ARXIV
  8. Operationalising the Superficial Alignment Hypothesis via Task Complexity 2026-02-18 · ARXIV
  9. Dex4D: Task-Agnostic Point Track Policy for Sim-to-Real Dexterous Manipulation 2026-02-18 · ARXIV
  10. Developing AI Agents with Simulated Data: Why, what, and how? 2026-02-18 · ARXIV
  11. CrispEdit: Low-Curvature Projections for Scalable Non-Destructive LLM Editing 2026-02-18 · ARXIV
  12. Avey-B 2026-02-18 · ARXIV
  13. Text Style Transfer with Parameter-efficient LLM Finetuning and Round-trip Translation 2026-02-17 · ARXIV
  14. Symmetry in language statistics shapes the geometry of model representations 2026-02-17 · ARXIV
  15. Scaling Beyond Masked Diffusion Language Models 2026-02-17 · ARXIV
  16. Rethinking Diffusion Models with Symmetries through Canonicalization with Applications to Molecular Graph Generation 2026-02-17 · ARXIV
  17. Long Context, Less Focus: A Scaling Gap in LLMs Revealed through Privacy and Personalization 2026-02-17 · ARXIV
  18. Hunt Globally: Deep Research AI Agents for Drug Asset Scouting in Investing, Business Development, and Search & Evaluation 2026-02-17 · ARXIV
  19. Efficient Sampling with Discrete Diffusion Models: Sharp and Adaptive Guarantees 2026-02-17 · ARXIV
  20. Cold-Start Personalization via Training-Free Priors from Structured World Models 2026-02-17 · ARXIV
  21. BPP: Long-Context Robot Imitation Learning by Focusing on Key History Frames 2026-02-17 · ARXIV
  22. When Benchmarks Lie: Evaluating Malicious Prompt Classifiers Under True Distribution Shift 2026-02-17 · ARXIV
  23. UniWeTok: An Unified Binary Tokenizer with Codebook Size $\mathit{2^{128}}$ for Unified Multimodal Large Language Model 2026-02-17 · ARXIV
  24. Process-Supervised Multi-Agent Reinforcement Learning for Reliable Clinical Reasoning 2026-02-17 · ARXIV
  25. Knowing When Not to Answer: Abstention-Aware Scientific Reasoning 2026-02-17 · ARXIV
  26. Index Light, Reason Deep: Deferred Visual Ingestion for Visual-Dense Document Question Answering 2026-02-17 · ARXIV
  27. GPT-5 vs Other LLMs in Long Short-Context Performance 2026-02-17 · ARXIV
  28. Deep Dense Exploration for LLM Reinforcement Learning via Pivot-Driven Resampling 2026-02-17 · ARXIV
  29. Semantic Chunking and the Entropy of Natural Language 2026-02-16 · ARXIV
  30. Realistic Face Reconstruction from Facial Embeddings via Diffusion Models 2026-02-16 · ARXIV
  31. In-Context Autonomous Network Incident Response: An End-to-End Large Language Model Agent Approach 2026-02-16 · ARXIV
  32. Improved Regret Guarantees for Online Mirror Descent using a Portfolio of Mirror Maps 2026-02-16 · ARXIV
  33. Imitating What Works: Simulation-Filtered Modular Policy Learning from Human Videos 2026-02-16 · ARXIV
  34. CoPE-VideoLM: Codec Primitives For Efficient Video Language Models 2026-02-16 · ARXIV
  35. Asynchronous Verified Semantic Caching for Tiered LLM Architectures 2026-02-16 · ARXIV
  36. Creative Ownership in the Age of AI 2026-02-14 · ARXIV
  37. UniT: Unified Multimodal Chain-of-Thought Test-time Scaling 2026-02-13 · ARXIV
  38. Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action Alignment 2026-02-13 · ARXIV
  39. On-Policy Context Distillation for Language Models 2026-02-13 · ARXIV
  40. MonarchRT: Efficient Attention for Real-Time Video Generation 2026-02-13 · ARXIV
  41. CM2: Reinforcement Learning with Checklist Rewards for Multi-Turn and Multi-Step Agentic Tool Use 2026-02-13 · ARXIV
  42. AttentionRetriever: Attention Layers are Secretly Long Document Retrievers 2026-02-13 · ARXIV
  43. Agentic Test-Time Scaling for WebAgents 2026-02-13 · ARXIV
  44. The Pensieve Paradigm: Stateful Language Models Mastering Their Own Context 2026-02-13 · ARXIV
  45. Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated Penalty 2026-02-13 · ARXIV
  46. P-GenRM: Personalized Generative Reward Model with Test-time User-based Scaling 2026-02-13 · ARXIV
  47. On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial Coverage 2026-02-13 · ARXIV
  48. Meta-Sel: Efficient Demonstration Selection for In-Context Learning via Supervised Meta-Learning 2026-02-13 · ARXIV
  49. Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation 2026-02-13 · ARXIV
  50. KAN-FIF: Spline-Parameterized Lightweight Physics-based Tropical Cyclone Estimation on Meteorological Satellite 2026-02-13 · ARXIV