03 / ARTICLE LEDGER

关联文章

按发布时间倒序

  1. OrLog: Resolving Complex Queries with LLMs and Probabilistic Reasoning 2026-02-02 · ARXIV
  2. From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching 2026-02-02 · ARXIV
  3. Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures 2026-02-02 · ARXIV
  4. Chain-of-thought obfuscation learned from output supervision can generalise to unseen tasks 2026-02-02 · ARXIV
  5. CATTO: Balancing Preferences and Confidence in Language Models 2026-02-02 · ARXIV
  6. #490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast 2026-02-01 · BLOGS_PODCASTS
  7. langbot-app/LangBot 2026-01-31 · GITHUB_TRENDING
  8. UEval: A Benchmark for Unified Multimodal Generation 2026-01-30 · ARXIV
  9. RedSage: A Cybersecurity Generalist LLM 2026-01-30 · ARXIV
  10. Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive Inquirers 2026-01-30 · ARXIV
  11. FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale 2026-01-30 · ARXIV
  12. DynaWeb: Model-Based Reinforcement Learning of Web Agents 2026-01-30 · ARXIV
  13. Temporal Guidance for Large Language Models 2026-01-30 · ARXIV
  14. Language-based Trial and Error Falls Behind in the Era of Experience 2026-01-30 · ARXIV
  15. EWSJF: An Adaptive Scheduler with Hybrid Partitioning for Mixed-Workload LLM Inference 2026-01-30 · ARXIV
  16. Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems 2026-01-30 · ARXIV
  17. When Flores Bloomz Wrong: Cross-Direction Contamination in Machine Translation Evaluation 2026-01-29 · ARXIV
  18. SokoBench: Evaluating Long-Horizon Planning and Reasoning in Large Language Models 2026-01-29 · ARXIV
  19. Evolutionary Strategies lead to Catastrophic Forgetting in LLMs 2026-01-29 · ARXIV
  20. Deep Researcher with Sequential Plan Reflection and Candidates Crossover (Deep Researcher Reflect Evolve) 2026-01-29 · ARXIV
  21. jeecgboot/JeecgBoot 2026-01-29 · GITHUB_TRENDING
  22. Show HN: A MitM proxy to see what your LLM tools are sending 2026-01-29 · HACKER_NEWS
  23. Reflective Translation: Improving Low-Resource Machine Translation via Structured Self-Reflection 2026-01-28 · ARXIV
  24. Post-LayerNorm Is Back: Stable, ExpressivE, and Deep 2026-01-28 · ARXIV
  25. Evaluation of Oncotimia: An LLM based system for supporting tumour boards 2026-01-28 · ARXIV
  26. Calibration without Ground Truth 2026-01-28 · ARXIV
  27. Show HN: I wrapped the Zorks with an LLM 2026-01-28 · HACKER_NEWS
  28. Unsupervised Text Segmentation via Kernel Change-Point Detection on Sentence Embeddings 2026-01-27 · ARXIV
  29. Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability 2026-01-27 · ARXIV
  30. Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy Prefixes 2026-01-27 · ARXIV
  31. POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration 2026-01-27 · ARXIV
  32. MortalMATH: Evaluating the Conflict Between Reasoning Objectives and Emergency Contexts 2026-01-27 · ARXIV
  33. Design Techniques for LLM-Powered Interactive Storytelling: A Case Study of the Dramamancer System 2026-01-27 · ARXIV
  34. ctELM: Decoding and Manipulating Embeddings of Clinical Trials with Embedding Language Models 2026-01-27 · ARXIV
  35. Strategies for Span Labeling with Large Language Models 2026-01-26 · ARXIV
  36. Spatial-Agent: Agentic Geo-spatial Reasoning with Scientific Core Concepts 2026-01-26 · ARXIV
  37. Empowering Medical Equipment Sustainability in Low-Resource Settings: An AI-Powered Diagnostic and Support Platform for Biomedical Technicians 2026-01-26 · ARXIV
  38. AgentDrive: An Open Benchmark Dataset for Agentic AI Reasoning with LLM-Generated Scenarios in Autonomous Systems 2026-01-26 · ARXIV
  39. A Scalable Measure of Loss Landscape Curvature for Analyzing the Training Dynamics of LLMs 2026-01-26 · ARXIV
  40. Show HN: Only 1 LLM can fly a drone 2026-01-26 · HACKER_NEWS
  41. Unrolling the Codex agent loop 2026-01-25 · BLOGS_PODCASTS
  42. Structured Hints for Sample-Efficient Lean Theorem Proving 2026-01-25 · ARXIV
  43. Provable Robustness in Multimodal Large Language Models via Feature Space Smoothing 2026-01-25 · ARXIV
  44. LLM-in-Sandbox Elicits General Agentic Intelligence 2026-01-25 · ARXIV
  45. Learning to Discover at Test Time 2026-01-25 · ARXIV
  46. Challenges and Research Directions for Large Language Model Inference Hardware 2026-01-25 · HACKER_NEWS