03 / ARTICLE LEDGER
关联文章
按发布时间倒序
- Tim Gowers: What sort of maths are LLMs good at? 2026-08-12 · HACKER_NEWS
- [AINews] How to steal a Reasoning Trace 2026-08-12 · BLOGS_PODCASTS
- llama.cpp 2026-08-12 · HACKER_NEWS
- Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders 2026-08-12 · ARXIV
- ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls 2026-08-12 · ARXIV
- Stealing Reasoning Traces from Proprietary LLM APIs 2026-08-12 · ARXIV
- 🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery 2026-08-12 · BLOGS_PODCASTS
- Accelerate cyber defense with OpenAI and AWS: Daybreak Red & Daybreak Blue now available to eligible customers on Amazon Bedrock 2026-08-12 · BLOGS_PODCASTS
- Deploying Anthropic Claude apps gateway for AWS for enterprise workloads 2026-08-12 · BLOGS_PODCASTS
- How ONESTRUCTION built the Ishigaki-IDS foundation model with AWS GenAIIC 2026-08-12 · BLOGS_PODCASTS
- SHE: Trajectory-driven Safety Harness Evolution for LLM Agents 2026-08-12 · ARXIV
- Stealing Reasoning Traces from Proprietary LLM APIs 2026-08-12 · HACKER_NEWS
- How Pixieset achieved 35% AI feature adoption by solving the right problem with Amazon Bedrock 2026-08-12 · BLOGS_PODCASTS
- Consilience for Verifier-Free Test-Time Scaling 2026-08-11 · ARXIV
- Apple Silicon and macOS VMs: 11–16× Faster LLM Inference with Llama.cpp 2026-08-11 · HACKER_NEWS
- How to organize Claude Code for product work 2026-08-11 · HACKER_NEWS
- Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness 2026-08-11 · ARXIV
- From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch 2026-08-11 · ARXIV
- How Claude marks AI-generated content 2026-08-11 · HACKER_NEWS
- [AINews] Muse Glimmer and Spark: Open Weights return Personal Superintelligence promise 2026-08-11 · BLOGS_PODCASTS
- Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions 2026-08-11 · ARXIV
- KGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMs 2026-08-11 · ARXIV
- A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy 2026-08-11 · ARXIV
- Humanising LLM Outputs Is Dumb 2026-08-11 · HACKER_NEWS
- Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots 2026-08-11 · HACKER_NEWS
- Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits 2026-08-11 · ARXIV
- Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing 2026-08-11 · ARXIV
- PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents 2026-08-11 · ARXIV
- An Exploratory Evaluation of LLM-Assisted Rewriting of Moderate-Complexity Financial Sentences for DisCoCat-Based Sentiment Analysis 2026-08-10 · ARXIV
- Model ML completes finance work more efficiently with GPT-5.6 Sol 2026-08-10 · BLOGS_PODCASTS
- Mistral Patent for "Code implemented tool calls" 2026-08-10 · HACKER_NEWS
- Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools 2026-08-10 · ARXIV
- SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent 2026-08-10 · ARXIV
- CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity 2026-08-10 · ARXIV
- Auto mode is now the default in Claude Code 2026-08-10 · HACKER_NEWS
- 我的 AI 工作流写了两年,直到 Opus 4.8 才真正生效 2026-08-10 · JUEJIN
- How I use LLMs to learn complex topics 2026-08-10 · HACKER_NEWS
- Building a Rust Inference Engine That Matches Llama.cpp 2026-08-09 · HACKER_NEWS
- The CPU is back: Rethinking the CPU-GPU split for LLM inference 2026-08-08 · HACKER_NEWS
- 开源项目第180期:Omnigent — Databricks 出品的 AI Agent 元编排层,让 Claude Code、Codex、Cursor 统一管控 2026-08-08 · JUEJIN
- [AINews] Zawinski's Law of MultiAgents 2026-08-08 · BLOGS_PODCASTS
- RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction 2026-08-08 · ARXIV
- The Claudyssey: A line-for-line translation of Homer's Odyssey by Claude Fable 5 2026-08-08 · HACKER_NEWS
- Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents 2026-08-08 · ARXIV
- Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents 2026-08-08 · ARXIV
- Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data 2026-08-08 · ARXIV
- How TReNDS automates root-cause analysis with Amazon Bedrock 2026-08-08 · BLOGS_PODCASTS
- TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories 2026-08-08 · ARXIV
- RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer 2026-08-08 · ARXIV
- The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2026-08-07 · ARXIV