03 / ARTICLE LEDGER
关联文章
按发布时间倒序
- Subliminal Effects in Your Data: A General Mechanism via Log-Linearity 2026-02-05 · ARXIV
- Rethinking the Trust Region in LLM Reinforcement Learning 2026-02-05 · ARXIV
- Reinforced Attention Learning 2026-02-05 · ARXIV
- Multi-Head LatentMoE and Head Parallel: Communication-Efficient and Deterministic MoE Parallelism 2026-02-05 · ARXIV
- CoT is Not the Chain of Truth: An Empirical Internal Analysis of Reasoning LLMs for Fake News Generation 2026-02-05 · ARXIV
- Understanding and Exploiting Weight Update Sparsity for Communication-Efficient Distributed RL 2026-02-04 · ARXIV
- Accelerating Scientific Research with Gemini: Case Studies and Common Techniques 2026-02-04 · ARXIV
- Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability 2026-02-03 · ARXIV
- RLAnything: Forge Environment, Policy, and Reward Model in Completely Dynamic RL System 2026-02-03 · ARXIV
- Reward-free Alignment for Conflicting Objectives 2026-02-03 · ARXIV
- RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents 2026-02-03 · ARXIV
- MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents 2026-02-03 · ARXIV
- Expanding the Capabilities of Reinforcement Learning via Text Feedback 2026-02-03 · ARXIV
- AgentRx: Diagnosing AI Agent Failures from Execution Trajectories 2026-02-03 · ARXIV
- Show HN: I built "AI Wattpad" to eval LLMs on fiction 2026-02-03 · HACKER_NEWS
- Sparse Reward Subsystem in Large Language Models 2026-02-03 · ARXIV
- Reliable Use of Lemmas via Eligibility Reasoning and Section$-$Aware Reinforcement Learning 2026-02-03 · ARXIV
- Reasoning and Tool-use Compete in Agentic RL:From Quantifying Interference to Disentangled Tuning 2026-02-03 · ARXIV
- How RLHF Amplifies Sycophancy 2026-02-03 · ARXIV
- Error Taxonomy-Guided Prompt Optimization 2026-02-03 · ARXIV
- UPA: Unsupervised Prompt Agent via Tree-Based Search and Selection 2026-02-02 · ARXIV
- TEON: Tensorized Orthonormalization Beyond Layer-Wise Muon for Large Language Model Pre-Training 2026-02-02 · ARXIV
- FOCUS: DLLMs Know How to Tame Their Compute Bound 2026-02-02 · ARXIV
- My iPhone 16 Pro Max produces garbage output when running MLX LLMs 2026-02-02 · HACKER_NEWS
- Safer Policy Compliance with Dynamic Epistemic Fallback 2026-02-02 · ARXIV
- OrLog: Resolving Complex Queries with LLMs and Probabilistic Reasoning 2026-02-02 · ARXIV
- From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching 2026-02-02 · ARXIV
- Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures 2026-02-02 · ARXIV
- Chain-of-thought obfuscation learned from output supervision can generalise to unseen tasks 2026-02-02 · ARXIV
- CATTO: Balancing Preferences and Confidence in Language Models 2026-02-02 · ARXIV
- #490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast 2026-02-01 · BLOGS_PODCASTS
- UEval: A Benchmark for Unified Multimodal Generation 2026-01-30 · ARXIV
- RedSage: A Cybersecurity Generalist LLM 2026-01-30 · ARXIV
- Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive Inquirers 2026-01-30 · ARXIV
- FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale 2026-01-30 · ARXIV
- DynaWeb: Model-Based Reinforcement Learning of Web Agents 2026-01-30 · ARXIV
- Temporal Guidance for Large Language Models 2026-01-30 · ARXIV
- Language-based Trial and Error Falls Behind in the Era of Experience 2026-01-30 · ARXIV
- EWSJF: An Adaptive Scheduler with Hybrid Partitioning for Mixed-Workload LLM Inference 2026-01-30 · ARXIV
- Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems 2026-01-30 · ARXIV
- When Flores Bloomz Wrong: Cross-Direction Contamination in Machine Translation Evaluation 2026-01-29 · ARXIV
- SokoBench: Evaluating Long-Horizon Planning and Reasoning in Large Language Models 2026-01-29 · ARXIV
- Evolutionary Strategies lead to Catastrophic Forgetting in LLMs 2026-01-29 · ARXIV
- Deep Researcher with Sequential Plan Reflection and Candidates Crossover (Deep Researcher Reflect Evolve) 2026-01-29 · ARXIV
- Show HN: A MitM proxy to see what your LLM tools are sending 2026-01-29 · HACKER_NEWS
- Reflective Translation: Improving Low-Resource Machine Translation via Structured Self-Reflection 2026-01-28 · ARXIV
- Post-LayerNorm Is Back: Stable, ExpressivE, and Deep 2026-01-28 · ARXIV
- Evaluation of Oncotimia: An LLM based system for supporting tumour boards 2026-01-28 · ARXIV
- Calibration without Ground Truth 2026-01-28 · ARXIV
- Show HN: I wrapped the Zorks with an LLM 2026-01-28 · HACKER_NEWS