03 / ARTICLE LEDGER

关联文章

按发布时间倒序

  1. Tim Gowers: What sort of maths are LLMs good at? 2026-08-12 · HACKER_NEWS
  2. [AINews] How to steal a Reasoning Trace 2026-08-12 · BLOGS_PODCASTS
  3. llama.cpp 2026-08-12 · HACKER_NEWS
  4. Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders 2026-08-12 · ARXIV
  5. ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls 2026-08-12 · ARXIV
  6. Stealing Reasoning Traces from Proprietary LLM APIs 2026-08-12 · ARXIV
  7. 🔬The BioAI Phase Shift - Matthew McPartlon & Neil Patil, Chai Discovery 2026-08-12 · BLOGS_PODCASTS
  8. Accelerate cyber defense with OpenAI and AWS: Daybreak Red & Daybreak Blue now available to eligible customers on Amazon Bedrock 2026-08-12 · BLOGS_PODCASTS
  9. Deploying Anthropic Claude apps gateway for AWS for enterprise workloads 2026-08-12 · BLOGS_PODCASTS
  10. How ONESTRUCTION built the Ishigaki-IDS foundation model with AWS GenAIIC 2026-08-12 · BLOGS_PODCASTS
  11. SHE: Trajectory-driven Safety Harness Evolution for LLM Agents 2026-08-12 · ARXIV
  12. Stealing Reasoning Traces from Proprietary LLM APIs 2026-08-12 · HACKER_NEWS
  13. How Pixieset achieved 35% AI feature adoption by solving the right problem with Amazon Bedrock 2026-08-12 · BLOGS_PODCASTS
  14. Consilience for Verifier-Free Test-Time Scaling 2026-08-11 · ARXIV
  15. Apple Silicon and macOS VMs: 11–16× Faster LLM Inference with Llama.cpp 2026-08-11 · HACKER_NEWS
  16. How to organize Claude Code for product work 2026-08-11 · HACKER_NEWS
  17. Decoding-Level Taboo: A Diagnostic Stress Test for LLM Robustness 2026-08-11 · ARXIV
  18. From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch 2026-08-11 · ARXIV
  19. How Claude marks AI-generated content 2026-08-11 · HACKER_NEWS
  20. [AINews] Muse Glimmer and Spark: Open Weights return Personal Superintelligence promise 2026-08-11 · BLOGS_PODCASTS
  21. Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions 2026-08-11 · ARXIV
  22. KGCaRe: Explainable Complex Conditional Question Answering using Automatic Knowledge Graph Construction and Context Retrieval with LLMs 2026-08-11 · ARXIV
  23. A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy 2026-08-11 · ARXIV
  24. Humanising LLM Outputs Is Dumb 2026-08-11 · HACKER_NEWS
  25. Show HN: Needle2: 14MB agentic LLM for phones, wearables, smart home and robots 2026-08-11 · HACKER_NEWS
  26. Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits 2026-08-11 · ARXIV
  27. Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing 2026-08-11 · ARXIV
  28. PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents 2026-08-11 · ARXIV
  29. An Exploratory Evaluation of LLM-Assisted Rewriting of Moderate-Complexity Financial Sentences for DisCoCat-Based Sentiment Analysis 2026-08-10 · ARXIV
  30. Model ML completes finance work more efficiently with GPT-5.6 Sol 2026-08-10 · BLOGS_PODCASTS
  31. Mistral Patent for "Code implemented tool calls" 2026-08-10 · HACKER_NEWS
  32. Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools 2026-08-10 · ARXIV
  33. SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent 2026-08-10 · ARXIV
  34. CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity 2026-08-10 · ARXIV
  35. Auto mode is now the default in Claude Code 2026-08-10 · HACKER_NEWS
  36. 我的 AI 工作流写了两年,直到 Opus 4.8 才真正生效 2026-08-10 · JUEJIN
  37. How I use LLMs to learn complex topics 2026-08-10 · HACKER_NEWS
  38. Building a Rust Inference Engine That Matches Llama.cpp 2026-08-09 · HACKER_NEWS
  39. The CPU is back: Rethinking the CPU-GPU split for LLM inference 2026-08-08 · HACKER_NEWS
  40. 开源项目第180期:Omnigent — Databricks 出品的 AI Agent 元编排层,让 Claude Code、Codex、Cursor 统一管控 2026-08-08 · JUEJIN
  41. [AINews] Zawinski's Law of MultiAgents 2026-08-08 · BLOGS_PODCASTS
  42. RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction 2026-08-08 · ARXIV
  43. The Claudyssey: A line-for-line translation of Homer's Odyssey by Claude Fable 5 2026-08-08 · HACKER_NEWS
  44. Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard Documents 2026-08-08 · ARXIV
  45. Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational Agents 2026-08-08 · ARXIV
  46. Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational Data 2026-08-08 · ARXIV
  47. How TReNDS automates root-cause analysis with Amazon Bedrock 2026-08-08 · BLOGS_PODCASTS
  48. TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent Trajectories 2026-08-08 · ARXIV
  49. RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning Transfer 2026-08-08 · ARXIV
  50. The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping 2026-08-07 · ARXIV