03 / ARTICLE LEDGER

关联文章

按发布时间倒序

  1. Subliminal Effects in Your Data: A General Mechanism via Log-Linearity 2026-02-05 · ARXIV
  2. Rethinking the Trust Region in LLM Reinforcement Learning 2026-02-05 · ARXIV
  3. Reinforced Attention Learning 2026-02-05 · ARXIV
  4. Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism 2026-02-05 · ARXIV
  5. CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation 2026-02-05 · ARXIV
  6. Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL 2026-02-04 · ARXIV
  7. Accelerating Scientific Research with Gemini: Case Studies and Common Techniques 2026-02-04 · ARXIV
  8. Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability 2026-02-03 · ARXIV
  9. RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System 2026-02-03 · ARXIV
  10. Reward-free Alignment for Conflicting Objectives 2026-02-03 · ARXIV
  11. RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents 2026-02-03 · ARXIV
  12. MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents 2026-02-03 · ARXIV
  13. Expanding the Capabilities of Reinforcement Learning via Text Feedback 2026-02-03 · ARXIV
  14. AgentRx: Diagnosing AI Agent Failures from Execution Trajectories 2026-02-03 · ARXIV
  15. Show HN: I built "AI Wattpad" to eval LLMs on fiction 2026-02-03 · HACKER_NEWS
  16. Sparse Reward Subsystem in Large Language Models 2026-02-03 · ARXIV
  17. Reliable Use of Lemmas via Eligibility Reasoning and Section$-$Aware Reinforcement Learning 2026-02-03 · ARXIV
  18. Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning 2026-02-03 · ARXIV
  19. How RLHF Amplifies Sycophancy 2026-02-03 · ARXIV
  20. Error Taxonomy-Guided Prompt Optimization 2026-02-03 · ARXIV
  21. UPA: Unsupervised Prompt Agent via Tree-Based Search and Selection 2026-02-02 · ARXIV
  22. TEON: Tensorized Orthonormalization Beyond Layer-Wise Muon for Large Language Model Pre-Training 2026-02-02 · ARXIV
  23. FOCUS: DLLMs Know How to Tame Their Compute Bound 2026-02-02 · ARXIV
  24. My iPhone 16 Pro Max produces garbage output when running MLX LLMs 2026-02-02 · HACKER_NEWS
  25. Safer Policy Compliance with Dynamic Epistemic Fallback 2026-02-02 · ARXIV
  26. OrLog: Resolving Complex Queries with LLMs and Probabilistic Reasoning 2026-02-02 · ARXIV
  27. From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching 2026-02-02 · ARXIV
  28. Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures 2026-02-02 · ARXIV
  29. Chain-of-thought obfuscation learned from output supervision can generalise to unseen tasks 2026-02-02 · ARXIV
  30. CATTO: Balancing Preferences and Confidence in Language Models 2026-02-02 · ARXIV
  31. #490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast 2026-02-01 · BLOGS_PODCASTS
  32. UEval: A Benchmark for Unified Multimodal Generation 2026-01-30 · ARXIV
  33. RedSage: A Cybersecurity Generalist LLM 2026-01-30 · ARXIV
  34. Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive Inquirers 2026-01-30 · ARXIV
  35. FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale 2026-01-30 · ARXIV
  36. DynaWeb: Model-Based Reinforcement Learning of Web Agents 2026-01-30 · ARXIV
  37. Temporal Guidance for Large Language Models 2026-01-30 · ARXIV
  38. Language-based Trial and Error Falls Behind in the Era of Experience 2026-01-30 · ARXIV
  39. EWSJF: An Adaptive Scheduler with Hybrid Partitioning for Mixed-Workload LLM Inference 2026-01-30 · ARXIV
  40. Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems 2026-01-30 · ARXIV
  41. When Flores Bloomz Wrong: Cross-Direction Contamination in Machine Translation Evaluation 2026-01-29 · ARXIV
  42. SokoBench: Evaluating Long-Horizon Planning and Reasoning in Large Language Models 2026-01-29 · ARXIV
  43. Evolutionary Strategies lead to Catastrophic Forgetting in LLMs 2026-01-29 · ARXIV
  44. Deep Researcher with Sequential Plan Reflection and Candidates Crossover (Deep Researcher Reflect Evolve) 2026-01-29 · ARXIV
  45. Show HN: A MitM proxy to see what your LLM tools are sending 2026-01-29 · HACKER_NEWS
  46. Reflective Translation: Improving Low-Resource Machine Translation via Structured Self-Reflection 2026-01-28 · ARXIV
  47. Post-LayerNorm Is Back: Stable, ExpressivE, and Deep 2026-01-28 · ARXIV
  48. Evaluation of Oncotimia: An LLM based system for supporting tumour boards 2026-01-28 · ARXIV
  49. Calibration without Ground Truth 2026-01-28 · ARXIV
  50. Show HN: I wrapped the Zorks with an LLM 2026-01-28 · HACKER_NEWS