LLM 관련 주요 논문 - 2026-07-20

1. CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data


2. DSWorld: A Data Science World Model for Efficient Autonomous Agents


3. Knowledge-Centric Agents for Workflow Generation


4. AgentFAIR: A Multi-Agent Collaborative Framework for FAIRness Evaluation of Geospatial Datasets


5. NeurOWL: An LLM-Based Neural-symbolic Framework for Incomplete OWL Ontology Reasoning


6. Behavioral Controllability of Agentic Models for Information Extraction: From Fixed Workflows to Reflective Agents


7. S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation


8. ToolVerse: Unlocking Massive Environments and Long-Horizon Tasks for Agentic Reinforcement Learning


9. Neuro-Symbolic AI for LEED compliance: Document-Centric Benchmarking, Deterministic Numeric Checking, and When Multimodal Hurts


10. MGDT: MLLM-Guided Diffusion Transformer with Relation-Adaptive Mixture-of-Experts for Multimodal Knowledge Graph Completion


11. Beyond a Joke: Multi-Angle Reasoning for Detecting and Explaining Harmful Humor in Memes


12. DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings


13. AnovaX: A Local, Multi-Agent Voice Assistant with LLM Planning, Typed Executors, and Adaptive Recovery


14. Cura 1T: Specialized Model for Agentic Healthcare


15. Causal-Audit: Explicit and Auditable Graph-based Reasoning via Target-Aware Causal Chain Construction


16. GraphDx: A Cost-Aware Knowledge-Enhanced Multi-Agent Framework for Sequential Diagnosis


17. Evaluating Open-Weight LLMs for Generating Structured Threat Information for Autonomous Vehicle Vulnerabilities


18. An Exam for Active Observers


19. When Do Multi-Agent Systems Help? An Information Bottleneck Perspective


20. Understanding Reasoning from Pretraining to Post-Training


21. LLM-Powered Agentic AI for 5G/6G Networks: A Tutorial and Survey on Architectures, Protocols, and Standardization


22. Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning


23. Induction in Both Directions: A Mechanistic Analysis of In-Context Learning in Masked Diffusion Language Models


24. Perceived AGI: Believability as Dimensional Completeness, Not Capability


25. CAMMAR: Culture-Aware Matryoshka for Metaphorical Arabic Representations


26. Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding


27. Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling


28. IMBench: A Benchmark for Intuitive Robotic Manipulation


29. Think at 5 Hz, Act at 20 Hz: Asynchronous Fast-Slow Vision-Language-Action Inference for Closed-Loop Driving


30. AEGIS: Assay-Aware Protocol Validation and Runtime Monitoring for Open-Source Liquid Handling Robots


31. Process Reward Informed Tree Rollout for Effective Multi-Turn RL


32. Scalable LLM Agent Tool Access in the Cloud


33. Ask Twice, Look Twice: Prompt Echoing Resolves the Question-First Paradox in Vision-Language Models


34. Evolutionary Algorithm-Guided LLMs for Physics-Informed Neural Network Design


35. CoWeaver: A Bi-directional, Learnable and Explainable Matching Engine for Mixed Human-Agent Science Collaboration


36. Kolmogorov–Arnold Networks for Small Language Models


37. SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification


38. Cache-Aware Prompt Compression:A Two-Tier Cost Model for LLM API Caching


39. LLM-Driven AutoML for Cross-Lingual Handwritten OCR: Closed-Loop Neural Architecture Search with GPT-5, GPT-4o, and Claude Sonnet 4


40. Verbalizable Representations Form a Global Workspace in Language Models


41. Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation


42. AI Trading: Evaluating Large Language Models for Technical Market Analysis


43. Large Language Models as Unified Multimodal Learners for Clinical Prediction