13Scaling Verification Can Be More Effective than Scaling Policy Learning for Vision-Language-Action AlignmentARXIV ArXiv阅读文章 arrow_forward
13Dario Amodei – "We are near the end of the exponential" [video]HACKER_NEWS Hacker News阅读文章 arrow_forward
13Customize AI agent browsing with proxies, profiles, and extensions in Amazon Bedrock AgentCore Browser | Amazon Web ServicesBLOGS_PODCASTS 博客与播客 AI Agent阅读文章 arrow_forward
13CM2: Reinforcement Learning with Checklist Rewards for Multi-Turn and Multi-Step Agentic Tool UseARXIV ArXiv AI Agent阅读文章 arrow_forward
13AttentionRetriever: Attention Layers are Secretly Long Document RetrieversARXIV ArXiv RAG阅读文章 arrow_forward
13Show HN: Moltis – AI assistant with memory, tools, and self-extending skillsHACKER_NEWS Hacker News阅读文章 arrow_forward
13Introducing Lockdown Mode and Elevated Risk labels in ChatGPTBLOGS_PODCASTS 博客与播客 AI Agent阅读文章 arrow_forward
13Show HN: Skill that lets Claude Code/Codex spin up VMs and GPUsHACKER_NEWS Hacker News阅读文章 arrow_forward
13I ditched OpenClaw and built a more secure AI agent (Blink and Mac Mini)HACKER_NEWS Hacker News AI Agent阅读文章 arrow_forward
13New Nick Bostrom Paper: Optimal Timing for Superintelligence [pdf]HACKER_NEWS Hacker News阅读文章 arrow_forward
13Evaluating Multilingual, Context-Aware Guardrails: A Humanitarian LLM Use CaseHACKER_NEWS Hacker News 大语言模型阅读文章 arrow_forward
13The Pensieve Paradigm: Stateful Language Models Mastering Their Own ContextARXIV ArXiv AI Agent阅读文章 arrow_forward
13Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated PenaltyARXIV ArXiv阅读文章 arrow_forward
13P-GenRM: Personalized Generative Reward Model with Test-time User-based ScalingARXIV ArXiv 大语言模型阅读文章 arrow_forward
13On the Complexity of Offline Reinforcement Learning with $Q^\star$-Approximation and Partial CoverageARXIV ArXiv阅读文章 arrow_forward
13Meta-Sel: Efficient Demonstration Selection for In-Context Learning via Supervised Meta-LearningARXIV ArXiv 大语言模型阅读文章 arrow_forward
13Learning beyond Teacher: Generalized On-Policy Distillation with Reward ExtrapolationARXIV ArXiv阅读文章 arrow_forward
13KAN-FIF: Spline-Parameterized Lightweight Physics-based Tropical Cyclone Estimation on Meteorological SatelliteARXIV ArXiv阅读文章 arrow_forward
12TabICLv2: A better, faster, scalable, and open tabular foundation modelARXIV ArXiv阅读文章 arrow_forward
12SCRAPL: Scattering Transform with Random Paths for Machine LearningARXIV ArXiv 机器学习阅读文章 arrow_forward
12Data-Efficient Hierarchical Goal-Conditioned Reinforcement Learning via Normalizing FlowsARXIV ArXiv阅读文章 arrow_forward
12Data Repetition Beats Data Scaling in Long-CoT Supervised Fine-TuningARXIV ArXiv 大语言模型阅读文章 arrow_forward
12Build long-running MCP servers on Amazon Bedrock AgentCore with Strands Agents integration | Amazon Web ServicesBLOGS_PODCASTS 博客与播客 MCP阅读文章 arrow_forward
12AI meets HR: Transforming talent acquisition with Amazon Bedrock | Amazon Web ServicesBLOGS_PODCASTS 博客与播客 AI Agent阅读文章 arrow_forward
12Anthropic raises $30B in Series G funding at $380B post-money valuationHACKER_NEWS Hacker News阅读文章 arrow_forward
12Beginning fully autonomous operations with the 6th-generation Waymo driverHACKER_NEWS Hacker News阅读文章 arrow_forward
12Gemini 3 Deep Think: Advancing science, research and engineeringBLOGS_PODCASTS 博客与播客 AI Agent阅读文章 arrow_forward
12Improving 15 LLMs at Coding in One Afternoon. Only the Harness ChangedHACKER_NEWS Hacker News 大语言模型阅读文章 arrow_forward