LLM 관련 주요 논문 - 2026-07-09

1. SkillCenter: A Large-Scale Source-Grounded Skill Library for Autonomous AI Agents


2. Do LLM-Generated Skills Make Better AI Data Scientists? A Component Ablation Across Data-Science Workflows


3. SpaCellAgent: A Self-Evolving LLM-Based Multi-Agent Framework for Trajectory Analysis


4. The Blind Curator: How a Biased Judge Silently Disables Skill Retirement in Self-Evolving Agents


5. Reason Less, Verify More: Deterministic Gates Recover a Silent Policy-Violation Failure Mode in Tool-Using LLM Agents


6. MIRA-Math: A Benchmark for Minimal Information Requesting and Mathematical Reasoning


7. Physics-Audited Agentic Discovery in Scientific Machine Learning


8. From Atomic Actions to Standard Operating Procedures: Iterative Tool Optimization for Self-Evolving LLM Agents


9. Does AI Understand Imaging? A Systematic Benchmark of Agentic AI for Computational Imaging Tasks


10. Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety


11. Learning social norms enhances compatibility in dynamic human-AI coordination


12. Large Behavior Model: A Promptable Digital Twin of the Retail Customer


13. Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics


14. LLM-powered reasoning in agent-based modeling


15. When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning


16. AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation


17. Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning


18. Co-LMLM: Continuous-Query Limited Memory Language Models


19. Breaking Database Lock-in: Agentic Regeneration of High Performance Storage Readers for Database Bypass


20. DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation


21. Future Confidence Distillation in Large Language Models


22. CARLA-GS: Decoupling Representation, Reasoning, and Physics Simulation for Autonomous Driving Corner-Case Synthesis


23. Creativity from Friction: Human-AI Interaction for Exploratory Structural Design


24. Single-Rollout Asynchronous Optimization for Agentic Reinforcement Learning


25. HIVE: Understanding Post-Hallucination Reasoning in Vision Language Models


26. SynthAVE: Scalable Synthetic Labeling for E-Commerce with LLM-Arena Validation


27. Heterogeneity-Adaptive Diffusion Schrodinger Bridge for PET-Guided Whole-Body MRI Translation


28. When Prompts Ignore Structure: Graph-Based Attribute Reasoning for Calibrated VLMs


29. On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces


30. Predicting LLM Safety Before Release by Simulating Deployment


31. Entropy Pacing Policy Optimization for Multi-Task Agentic Reinforcement Learning


32. Tree-of-Thoughts Reasoning for Text-to-Image In-Context Learning


33. Progressive Crystallization: Turning Agent Exploration into Deterministic, Lower-Cost Workflows in Production


34. Riemannian Geometry for Pre-trained Language Model Embeddings


35. AnchorPrune: Relevance-Anchored Contextual Expansion for Visual Token Pruning


36. End-to-End LLM Flight Planning with RAG-based Memory and Multi-modal Coach Agent


37. Large Language Models (LLMs) and Generative AI in Cybersecurity and Privacy: A Survey of Dual-Use Risks, AI-Generated Malware, Explainability, and Defensive Strategies


38. Comprehensive Evaluation of Large Language Model Responses: A Multi-Factor Scoring System


39. GemNav: Discrete-Token Visual Robot Navigation using a Multimodal Large Language Model


40. A Gold-Standard Study of What Makes a Lightweight Game-Playing Agent Strong


41. Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs


42. Ad Headline Generation using Self-Critical Masked Language Model


43. When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems


44. A Multi-Analyst LLM Pipeline for Auditable Rule Discovery Across 68 Public Physiological Corpora


45. What Predicts Correctness in Text-to-SQL? A Selective-Prediction Study


46. SmartHomeSecure: Automated Detection and Repair of Smart Home Configuration Errors Using Large Language Models


47. Reliable and Developer-Aligned Evaluation of Agents for Software Engineering


48. Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering


49. Specification Grounding Drives Test Effectiveness for LLM Code


50. Open-Ended Scenario Reasoning for Specialist Model Adaptation


51. LLM-Guided Task-Semantic Field Factorization for Industrial Process Forecasting


52. SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts


53. TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation


54. When Agents Remember Too Much: Memory Poisoning Attacks on Large Language Model Agents