03 / ARTICLE LEDGER
关联文章
按发布时间倒序
- OrLog: Resolving Complex Queries with LLMs and Probabilistic Reasoning 2026-02-02 · ARXIV
- From Similarity to Vulnerability: Key Collision Attack on LLM Semantic Caching 2026-02-02 · ARXIV
- Character as a Latent Variable in Large Language Models: A Mechanistic Account of Emergent Misalignment and Conditional Safety Failures 2026-02-02 · ARXIV
- Chain-of-thought obfuscation learned from output supervision can generalise to unseen tasks 2026-02-02 · ARXIV
- CATTO: Balancing Preferences and Confidence in Language Models 2026-02-02 · ARXIV
- #490 – State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast 2026-02-01 · BLOGS_PODCASTS
- langbot-app/LangBot 2026-01-31 · GITHUB_TRENDING
- UEval: A Benchmark for Unified Multimodal Generation 2026-01-30 · ARXIV
- RedSage: A Cybersecurity Generalist LLM 2026-01-30 · ARXIV
- Reasoning While Asking: Transforming Reasoning Large Language Models from Passive Solvers to Proactive Inquirers 2026-01-30 · ARXIV
- FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale 2026-01-30 · ARXIV
- DynaWeb: Model-Based Reinforcement Learning of Web Agents 2026-01-30 · ARXIV
- Temporal Guidance for Large Language Models 2026-01-30 · ARXIV
- Language-based Trial and Error Falls Behind in the Era of Experience 2026-01-30 · ARXIV
- EWSJF: An Adaptive Scheduler with Hybrid Partitioning for Mixed-Workload LLM Inference 2026-01-30 · ARXIV
- Epistemic Context Learning: Building Trust the Right Way in LLM-Based Multi-Agent Systems 2026-01-30 · ARXIV
- When Flores Bloomz Wrong: Cross-Direction Contamination in Machine Translation Evaluation 2026-01-29 · ARXIV
- SokoBench: Evaluating Long-Horizon Planning and Reasoning in Large Language Models 2026-01-29 · ARXIV
- Evolutionary Strategies lead to Catastrophic Forgetting in LLMs 2026-01-29 · ARXIV
- Deep Researcher with Sequential Plan Reflection and Candidates Crossover (Deep Researcher Reflect Evolve) 2026-01-29 · ARXIV
- jeecgboot/JeecgBoot 2026-01-29 · GITHUB_TRENDING
- Show HN: A MitM proxy to see what your LLM tools are sending 2026-01-29 · HACKER_NEWS
- Reflective Translation: Improving Low-Resource Machine Translation via Structured Self-Reflection 2026-01-28 · ARXIV
- Post-LayerNorm Is Back: Stable, ExpressivE, and Deep 2026-01-28 · ARXIV
- Evaluation of Oncotimia: An LLM based system for supporting tumour boards 2026-01-28 · ARXIV
- Calibration without Ground Truth 2026-01-28 · ARXIV
- Show HN: I wrapped the Zorks with an LLM 2026-01-28 · HACKER_NEWS
- Unsupervised Text Segmentation via Kernel Change-Point Detection on Sentence Embeddings 2026-01-27 · ARXIV
- Teaching Models to Teach Themselves: Reasoning at the Edge of Learnability 2026-01-27 · ARXIV
- Reuse your FLOPs: Scaling RL on Hard Problems by Conditioning on Very Off-Policy Prefixes 2026-01-27 · ARXIV
- POPE: Learning to Reason on Hard Problems via Privileged On-Policy Exploration 2026-01-27 · ARXIV
- MortalMATH: Evaluating the Conflict Between Reasoning Objectives and Emergency Contexts 2026-01-27 · ARXIV
- Design Techniques for LLM-Powered Interactive Storytelling: A Case Study of the Dramamancer System 2026-01-27 · ARXIV
- ctELM: Decoding and Manipulating Embeddings of Clinical Trials with Embedding Language Models 2026-01-27 · ARXIV
- Strategies for Span Labeling with Large Language Models 2026-01-26 · ARXIV
- Spatial-Agent: Agentic Geo-spatial Reasoning with Scientific Core Concepts 2026-01-26 · ARXIV
- Empowering Medical Equipment Sustainability in Low-Resource Settings: An AI-Powered Diagnostic and Support Platform for Biomedical Technicians 2026-01-26 · ARXIV
- AgentDrive: An Open Benchmark Dataset for Agentic AI Reasoning with LLM-Generated Scenarios in Autonomous Systems 2026-01-26 · ARXIV
- A Scalable Measure of Loss Landscape Curvature for Analyzing the Training Dynamics of LLMs 2026-01-26 · ARXIV
- Show HN: Only 1 LLM can fly a drone 2026-01-26 · HACKER_NEWS
- Unrolling the Codex agent loop 2026-01-25 · BLOGS_PODCASTS
- Structured Hints for Sample-Efficient Lean Theorem Proving 2026-01-25 · ARXIV
- Provable Robustness in Multimodal Large Language Models via Feature Space Smoothing 2026-01-25 · ARXIV
- LLM-in-Sandbox Elicits General Agentic Intelligence 2026-01-25 · ARXIV
- Learning to Discover at Test Time 2026-01-25 · ARXIV
- Challenges and Research Directions for Large Language Model Inference Hardware 2026-01-25 · HACKER_NEWS