09Improving AI models’ ability to explain their predictionsBLOGS_PODCASTS 博客与播客 大语言模型阅读文章 arrow_forward
09Show HN: VS Code Agent Kanban: Task Management for the AI-Assisted DeveloperHACKER_NEWS Hacker News AI Agent阅读文章 arrow_forward
09Nvidia backs AI data center startup Nscale as it hits $14.6B valuationHACKER_NEWS Hacker News阅读文章 arrow_forward
09Show HN: Mcp2cli – One CLI for every API, 96-99% fewer tokens than native MCPHACKER_NEWS Hacker News MCP阅读文章 arrow_forward
09Grammarly is offering ‘expert’ AI reviews from famous dead and living writersHACKER_NEWS Hacker News阅读文章 arrow_forward
09Show HN: Reviving a 20-year-old puzzle game Chromatron with Ghidra and AIHACKER_NEWS Hacker News阅读文章 arrow_forward
08We should revisit literate programming in the agent eraHACKER_NEWS Hacker News AI Agent阅读文章 arrow_forward
08Agent Safehouse – macOS-native sandboxing for local agentsHACKER_NEWS Hacker News AI Agent阅读文章 arrow_forward
08Oracle may slash up to 30k jobs to fund AI data-centers as US banks retreatHACKER_NEWS Hacker News阅读文章 arrow_forward
08Phi-4-reasoning-vision and the lessons of training a multimodal reasoning modelHACKER_NEWS Hacker News阅读文章 arrow_forward
08SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via CIHACKER_NEWS Hacker News AI Agent阅读文章 arrow_forward
08Autoresearch: Agents researching on single-GPU nanochat training automaticallyHACKER_NEWS Hacker News AI Agent阅读文章 arrow_forward
07LLM Doesn't Write Correct Code. It Writes Plausible CodeHACKER_NEWS Hacker News 大语言模型阅读文章 arrow_forward
07Sarvam 105B, the first competitive Indian open source LLMHACKER_NEWS Hacker News 大语言模型阅读文章 arrow_forward
07LLMs work best when the user defines their acceptance criteria firstHACKER_NEWS Hacker News 大语言模型阅读文章 arrow_forward
06Towards Provably Unbiased LLM Judges via Bias-Bounded EvaluationARXIV ArXiv 大语言模型阅读文章 arrow_forward