10CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and DiversityARXIV 大语言模型 cs.CL阅读文章 arrow_forward
10MirrorWorld: Taming Video Diffusion Models for Mirror Reflection GenerationARXIV 生成式 AI 计算机视觉阅读文章 arrow_forward
10Show HN: A replayable A2A jury for tracing how agents influence decisionsHACKER_NEWS AI Agent Hacker News阅读文章 arrow_forward
10Human vs. AI – Diff-based line-level provenance for text under agentic editingHACKER_NEWS AI Hacker News阅读文章 arrow_forward
09Amazon circumvents Gilroy community vote for AI data centerHACKER_NEWS AI Hacker News阅读文章 arrow_forward
09Real-time MCP interceptor that blocks .env reads and dangerous commands agentsHACKER_NEWS AI Agent Hacker News阅读文章 arrow_forward
09Denmark Requires Oral Defenses for Students' Written Work to Counter AI CheatingHACKER_NEWS AI Hacker News阅读文章 arrow_forward
09Building a Rust Inference Engine That Matches Llama.cppHACKER_NEWS 大语言模型 Hacker News阅读文章 arrow_forward
09Timeline of the OpenAI accidental attack against Hugging FaceHACKER_NEWS AI Hacker News阅读文章 arrow_forward
08The CPU is back: Rethinking the CPU-GPU split for LLM inferenceHACKER_NEWS 大语言模型 Hacker News阅读文章 arrow_forward
08开源项目第180期:Omnigent — Databricks 出品的 AI Agent 元编排层,让 Claude Code、Codex、Cursor 统一管控JUEJIN 大语言模型 AI Agent阅读文章 arrow_forward
08Beyond Top-K: Replacing Black-Box Retrieval with Interpretable Agentic OperationsARXIV AI Agent cs.AI阅读文章 arrow_forward
08RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward ConstructionARXIV 大语言模型 cs.LG阅读文章 arrow_forward
08The Claudyssey: A line-for-line translation of Homer's Odyssey by Claude Fable 5HACKER_NEWS 大语言模型 Hacker News阅读文章 arrow_forward
08Does FLAIR super-resolution erase or hallucinate small white-matter lesions?ARXIV 计算机视觉 cs.CV阅读文章 arrow_forward
08Responding to the next frontier of critical cyber capabilitiesBLOGS_PODCASTS AI Security阅读文章 arrow_forward
08Benchmarking and Enhancing LLMs for Rule-Intensive Review of National Standard DocumentsARXIV 大语言模型 AI Agent阅读文章 arrow_forward
08Determining playoff clinching scenarios in the NHL using constraint programmingBLOGS_PODCASTS 生成式 AI 机器学习阅读文章 arrow_forward
08Benchmarking the Benchmarks: Evaluating Benchmarks for Conversational AgentsARXIV 大语言模型 AI Agent阅读文章 arrow_forward
08Tytan: Interactive Neurosymbolic Construction of Analytic Semantic Schemas from Relational DataARXIV 大语言模型 cs.DB阅读文章 arrow_forward
08How TReNDS automates root-cause analysis with Amazon BedrockBLOGS_PODCASTS 大语言模型 RAG阅读文章 arrow_forward
08TRAJDEBUG: Tracing Error Lifecycle to Identify Critical Failures in Long-Horizon Agent TrajectoriesARXIV 大语言模型 AI Agent阅读文章 arrow_forward
08How Cohere Health digitizes clinical policies using Amazon Bedrock AgentCoreBLOGS_PODCASTS AI Agent 生成式 AI阅读文章 arrow_forward
08RP-OPSD: Reasoning-Pivot-Guided On-Policy Self-Distillation for Multilingual Reasoning TransferARXIV 大语言模型 cs.CL阅读文章 arrow_forward
07Challenges in Evaluating Explanation Methods for Static and Evolving DataARXIV AI cs.AI阅读文章 arrow_forward
07Kitesurf: Agent-first browser that runs in V8 isolatesHACKER_NEWS AI Agent Hacker News阅读文章 arrow_forward
07CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal TasksARXIV AI Agent cs.LG阅读文章 arrow_forward