03 / ARTICLE LEDGER

关联文章

按发布时间倒序

  1. A Visual Dependence-Aware Framework for Multimodal Unsupervised Continual Post-Training 2026-08-28 · ARXIV
  2. VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning 2026-08-27 · ARXIV
  3. LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training 2026-08-26 · ARXIV
  4. Predicting Multiple Clinical Outcomes Related to Functional Recovery and Social Isolation Among Older Adults After Lower-Limb Fracture or Hip Replacement 2026-08-26 · ARXIV
  5. EG-ARSA: An Expert-Grounded Open Model for Visual Road Safety Auditing in Low-Resource Settings 2026-08-25 · ARXIV
  6. AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design 2026-08-14 · ARXIV
  7. Diagram-MMU: A Multi-Modal Benchmark for Scientific Diagrams 2026-08-14 · ARXIV
  8. A Neighborhood Attention Transformer Network for Enhanced 3D Segmentation of the Left Anterior Descending Artery 2026-08-14 · ARXIV
  9. Beyond Trial-and-Error: Agentic Optimization for Image-to-Video Adherence 2026-08-13 · ARXIV
  10. Class Activation Mapping in Explainable Computer Vision: A Method-Centered Review of CNN, Transformer, and Foundation-Model-Era Visual Explanations 2026-08-13 · ARXIV
  11. DreamFly: Causal Memory and Receding-Horizon Diffusion Planning for Aerial Vision-Language Navigation 2026-08-13 · ARXIV
  12. AVA-Encoder: Towards Agent-Native Video Representation Learning 2026-08-13 · ARXIV
  13. AlbumentationsX: One Augmentation Pipeline for Images and Related Annotations 2026-08-13 · ARXIV
  14. MultiModal Code-Switching: Interleaving Visual Objects into Language for Explicit Object-Level Alignment 2026-08-13 · ARXIV
  15. Test-Time Self-Evolving GUI Visual Grounding via Reflection-Guided On-Policy Self-Distillation 2026-08-12 · ARXIV
  16. Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains 2026-08-12 · ARXIV
  17. Financial Numerical Prediction and Allocation as Token Generation 2026-08-12 · ARXIV
  18. Space-Creating versus Dead Possession: An Off-Ball Possession-Quality Index for Broadcast Football 2026-08-12 · ARXIV
  19. Multimodal Model Diffing for Feature Discovery and Control 2026-08-11 · ARXIV
  20. SABRE: Scalable and Automated Benchmarking of VLMs under Stress 2026-08-11 · ARXIV
  21. MirrorWorld: Taming Video Diffusion Models for Mirror Reflection Generation 2026-08-10 · ARXIV
  22. Does FLAIR super-resolution erase or hallucinate small white-matter lesions? 2026-08-08 · ARXIV
  23. Predicting Brain Morphometry with MT-GNN: Mesh Evolution in Continuous Time with Graph-Based Metric Tensor Embeddings 2026-08-07 · ARXIV
  24. ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs 2026-08-05 · ARXIV
  25. UEmbed: Unified Sparse and Dense Multimodal Embeddings 2026-08-04 · ARXIV
  26. VetClaw: An Edge-Cloud Multimodal Agentic System for Veterinary Disease Screening 2026-07-29 · ARXIV
  27. KANEx: Translating Kolmogorov-Arnold Networks' Interpretability to Medical Explainability 2026-07-29 · ARXIV
  28. Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation 2026-07-28 · ARXIV
  29. ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding 2026-07-28 · ARXIV
  30. SM4RT: Learning Structured Motion Geometry for 4D Reconstruction 2026-07-27 · ARXIV
  31. MIT researchers teach AI models to interpret charts 2026-06-03 · BLOGS_PODCASTS
  32. IVGT: Implicit Visual Geometry Transformer for Neural Scene Representation 2026-05-18 · ARXIV
  33. Seeing Fast and Slow: Learning the Flow of Time in Videos 2026-04-24 · ARXIV
  34. Generative AI improves a wireless vision system that sees through obstructions 2026-03-19 · BLOGS_PODCASTS
  35. Does AI See like Art Historians? Interpreting How Vision Language Models Recognize Artistic Style 2026-03-12 · ARXIV
  36. Improving AI models’ ability to explain their predictions 2026-03-09 · BLOGS_PODCASTS
  37. AI 智能体如何重构开发工作流 2026-03-08 · JUEJIN
  38. AI视觉连载8:传统 CV 之边缘检测 2026-03-03 · JUEJIN
  39. 拒绝模糊:用“空洞卷积”重塑深度学习的视野 2026-02-26 · JUEJIN
  40. AI 视觉连载5:传统 CV 之均值滤波 2026-02-24 · JUEJIN
  41. Build an intelligent photo search using Amazon Rekognition, Amazon Neptune, and Amazon Bedrock | Amazon Web Services 2026-02-24 · BLOGS_PODCASTS
  42. Accelerating AI model production at Hexagon with Amazon SageMaker HyperPod | Amazon Web Services 2026-02-23 · BLOGS_PODCASTS
  43. SCRAPL: Scattering Transform with Random Paths for Machine Learning 2026-02-12 · ARXIV
  44. AI 视觉连载3:RGB与通道 2026-02-11 · JUEJIN
  45. d2l-ai/d2l-zh 2026-02-05 · GITHUB_TRENDING
  46. MaaAssistantArknights/MaaAssistantArknights 2026-01-26 · GITHUB_TRENDING