26CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language ModelsARXIV ArXiv 大语言模型阅读文章 arrow_forward
26A Diversity Diet for a Healthier Model: A Case Study of French ModernBERTARXIV ArXiv 自然语言处理阅读文章 arrow_forward
25Efficiently serve dozens of fine-tuned models with vLLM on Amazon SageMaker AI and Amazon Bedrock | Amazon Web ServicesBLOGS_PODCASTS 博客与播客阅读文章 arrow_forward
25Why Pass@k Optimization Can Degrade Pass@1: Prompt Interference in LLM Post-trainingARXIV ArXiv 大语言模型阅读文章 arrow_forward
25Untied Ulysses: Memory-Efficient Context Parallelism via Headwise ChunkingARXIV ArXiv阅读文章 arrow_forward
25The Diffusion Duality, Chapter II: $Ψ$-Samplers and Efficient CurriculumARXIV ArXiv阅读文章 arrow_forward
25Learning from Trials and Errors: Reflective Test-Time Planning for Embodied LLMsARXIV ArXiv AI Agent阅读文章 arrow_forward
25PA bench: Evaluating web agents on real world personal assistant workflowsHACKER_NEWS Hacker News AI Agent阅读文章 arrow_forward
25Building intelligent event agents using Amazon Bedrock AgentCore and Amazon Bedrock Knowledge Bases | Amazon Web ServicesBLOGS_PODCASTS 博客与播客 RAG阅读文章 arrow_forward
25AI to help researchers see the bigger picture in cell biologyBLOGS_PODCASTS 博客与播客 机器学习阅读文章 arrow_forward
25Show HN: Sgai – Goal-driven multi-agent software dev (GOAL.md → working code)HACKER_NEWS Hacker News AI Agent阅读文章 arrow_forward
25Launch HN: TeamOut (YC W22) – AI agent for planning company retreatsHACKER_NEWS Hacker News AI Agent阅读文章 arrow_forward
25Show HN: A real-time strategy game that AI agents can playHACKER_NEWS Hacker News AI Agent阅读文章 arrow_forward
25Show HN: Context Mode – 315 KB of MCP output becomes 5.4 KB in Claude CodeHACKER_NEWS Hacker News MCP阅读文章 arrow_forward
25VAUQ: Vision-Aware Uncertainty Quantification for LVLM Self-EvaluationARXIV ArXiv 大语言模型阅读文章 arrow_forward
25Scaling Vision Transformers: Evaluating DeepSpeed for Image-Centric WorkloadsARXIV ArXiv阅读文章 arrow_forward
25ProxyFL: A Proxy-Guided Framework for Federated Semi-Supervised LearningARXIV ArXiv阅读文章 arrow_forward
25Localized Dynamics-Aware Domain Adaption for Off-Dynamics Offline Reinforcement LearningARXIV ArXiv阅读文章 arrow_forward
25Beyond the Star Rating: A Scalable Framework for Aspect-Based Sentiment Analysis Using LLMs and Text ClassificationARXIV ArXiv 大语言模型阅读文章 arrow_forward
25An Expert Schema for Evaluating Large Language Model Errors in Scholarly Question-Answering SystemsARXIV ArXiv 大语言模型阅读文章 arrow_forward
24Skill-Inject: Measuring Agent Vulnerability to Skill File AttacksARXIV ArXiv AI Agent阅读文章 arrow_forward
24Show HN: Moonshine Open-Weights STT models – higher accuracy than WhisperLargev3HACKER_NEWS Hacker News阅读文章 arrow_forward
24KNIGHT: Knowledge Graph-Driven Multiple-Choice Question Generation with Adaptive Hardness CalibrationARXIV ArXiv RAG阅读文章 arrow_forward
24JUCAL: Jointly Calibrating Aleatoric and Epistemic Uncertainty in Classification TasksARXIV ArXiv阅读文章 arrow_forward
24Behavior Learning (BL): Learning Hierarchical Optimization Structures from DataARXIV ArXiv 机器学习阅读文章 arrow_forward