LLM 관련 주요 논문 - 2026-07-29

1. CHARM: A Multimodal Graph Foundation Model with Hierarchical Context Modeling for Zero-Shot Transfer


2. Large Language Model for Operations Research Formulation Selection in Multi-Warehouse Inventory Allocation


3. A Cost-Effective Multimodal LLM Reasoning Framework for Question Answering over Irregular Clinical Time Series


4. Penelope: Localized Latent Recurrence for Efficient Structured Reasoning


5. Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks


6. HiSkill: Empowering LLM Agents with Hierarchical Skill Graphs


7. Speculate While You Reason: Teaching Agents to Predict Their Next Tool Call via Joint Agent-Speculator RL


8. Nudging Sustainable Choices through LLM-Generated Recommendation Explanations


9. Cognivia: A Cognitive Behavioral Therapy Copilot for Evidence-Based Mental Healthcare


10. DecoEvo: Score-Decoupled Co-Evolution of Solver and Rubric-Generator Skills in Text Space


11. OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs


12. CoRT: Counterfactual Replay for Token-Level Rubric-Guided Policy Optimization


13. Beyond Epistemia: Epistemic Schizologia and Large Language Models as Techno-Semiotic Machines


14. How Small Can You Go? A Controlled Study of LoRA Rank, Target Modules, and Quantization Trade-offs for Text-to-SQL on a 60M-Parameter Model


15. PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents


16. TRWH: A Text-Driven Random Walk Heterogeneous GNN for Semantic-Aware Sparse Recommendation


17. Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm


18. The Disruptive Impact of Large Language Models on Capture the Flag Competitions and the Path Toward Fair Play


19. A Control System, a Dataset, and a Recipe for Making Frozen LLM Agents Learn a Domain


20. Context Assembly as the Controlled Variable: A Control-Theoretic View of Harness Policies for Frozen LLM Agents


21. COVENANT: Natural-Language Workflow Compilation for Aligned Agent Execution


22. Cyber-Capable AI Agents: Vulnerabilities, Evaluation Containment, and Defensive Response


23. Cardiologent: Multi-Agent Clinical Decision Support for Patient-Level Arrhythmia Assessment, Urgency, and Management


24. Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe


25. ContractHIL-HLS: Contract-Aligned Multi-Agent Workflow with Hardware-in-the-Loop Feedback for HLS Design


26. The User Asks, Platforms Compete: How Agentic Recommendation Markets Take Shape


27. CADENCE: A Cardiac Atom Dictionary for Interpretable Neural Concept Extraction from ECG Foundation Models


28. Everyone is unique: Towards Behaviorally Heterogeneous Negotiation Dialogue Systems for Debt Collection


29. PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention


30. When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops


31. How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation


32. ScalableRAG: High-Quality RAG at Zero Ingestion Cost


33. Towards Robust Reinforcement Learning for Small-Scale Language Model Agents


34. Towards an Agent Operating System - Lessons from Classical and Cloud OS


35. How Often Should a Recommender Call an LLM? Value-Weighted Routing, Monitoring, and Seasonal Robustness


36. Addressable Recall Compaction for Long Context-Window Control in AI Agents


37. CogEEGAgent: Toward Autonomous Cognitive EEG Analysis with Grounded Execution and Selection-Aware Verification


38. Chart-Supported or Model-Supplied? Examining MLLM-Generated Claims for Accessible Visualization


39. MusiChat: Vibe Composing for Music Creation


40. AdaKP: Online Adaptive Knowledge-Point Selection for Reasoning-Oriented Reinforcement Learning


41. Aletheia: An Offline-First Clinical Decision Support System for Differential Diagnosis in Low-Resource Healthcare Settings


42. RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation


43. Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding


44. GLIDE: Guided Layerwise Hybrid Attention for Efficient LLM Inference


45. Unified Semantic Modeling Framework for Large-Scale Job Understanding at LinkedIn


46. Personalization, Personas, and Forecasting in Value Alignment


47. HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising


48. RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation


49. ProcAgent: An Agentic Framework for Procedural Task Guidance on Edge with Human-in-the-Loop


50. LLM Scheming Inversely Scales with Pretraining Language Coverage


51. PATHFinder Agent for Tailored Prenatal Care


52. Crystalis: Progressive Nucleation and Semantic Annealing for Coordinated Multi-View Visualization Generation


53. GrocLM: Grocery Category Recommendation in E-Commerce with Large Language Models


54. CaRE Compute-aware Remasking Evaluation Protocol for Masked Diffusion Language Models


55. Kernel Forge: An Agent Harness for LLM-based Generation and Optimization of CUDA Kernels


56. Beyond Memory: A Templated Substrate for Heterogeneous Collaborative Knowledge Work with LLM Agents


57. Do Models Fake Alignment Without Clear Consequences?


58. Does Runtime Topology Context Improve LLM-Generated Kubernetes Security Patches?


59. MemLens: A Value-Aware Memory Management System with Interactive Analytics for LLM-based Agents


60. Knowledge-Guided Multimodal Reasoning over Interacting Streams for Video-Level Ambivalence and Hesitancy Recognition


61. Detecting Knowledge Inconsistencies Across Text, Tables, and Knowledge Graphs


62. MODUS: Decoder-Only Any-to-Any Modeling of Diverse Modalities


63. Evaluating Multi-Turn Multimodal Diagnostic Reasoning on Challenging Real-World Clinical Cases


64. Evaluating VLMs for Autonomous Agent-Driven Geometry Clipping Detection in Video Game QA


65. Minimizing Targeted Activations: Input-Only Suppression of Evaluation-Awareness Latents in Large Language Models


66. Stemma: Induced Decision Regions Reveal LLM Provenance


67. How Do LLMs Read Bug Reports? An Empirical Study of Attention in LLMs for Automated Program Repair


68. OmniQEC: discovering practical quantum error-correcting codes by an AI scientist


69. Tools Are Not Islands: Set-Level Tool Retrieval for LLM Agents via Query-Conditioned Hyperedge Prediction


70. DynaBridge: Dynamic Summary-Guided Cross-Task Multimodal Fusion for DASS-Structured Mental Health Assessment


71. Why Public Service AI Governance Frameworks Risk Failing in the Age of General-Purpose AI: Lessons from Policing


72. KQFuzz: Knowledge-Guided Fuzzing for Quantum Libraries via Large Language Models


73. Construction-Driven Injection: Linguistically-Grounded Edit-Based Code-Mixing Fingerprints for Large Language Models


74. A Human-in-the-Loop Corpus for LLM-Based Simplification of Scientific Summaries


75. Beyond Self-Knowledge: Propagating Uncertainty Across Reasoning and Retrieval in LLMs


76. IRIS: Reusable Identity Representations from Frozen LLMs for Entity Alignment


77. Visual prompt engineering for video models


78. Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation


79. Architectural Backdoors in Vision-Language Model Supply Chains via Representation Steering


80. Emergent Latent-State Computation under Stochastic Volatility


81. MARS: Multi-Agent Re-ranking for Repeat-Order Food Delivery Recommendation


82. Every Time I Hire a Linguist, Inference Costs Go Down: On Linguistic Rules as Effective Prompt Compressors


83. Specula: Scaling formal specifications for autonomous model checking of system code


84. CAST: Game Solvers as Turn-Level Teachers for LLM Agents


85. Retraction-Free Optimization over the Stiefel Manifold for the LoRA Fine-Tuning


86. Hybrid Analysis for Secure MCP Tool Use in LLM Agents


87. Where Steering Signals Come From: Activation Source Selection in Activation Steering


88. Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks


89. VaLiDRec: Variable-Length LLM-Aligned Semantic IDs for Generative Recommendation


90. RIDGE: An Autonomous Framework for Validation and Method Discovery in LLM-Generated Option Pricing


91. TabRank: Chain-of-Thought Distillation for Table Re-Rankers


92. Agentic AI for Scientific Reasoning in Autonomous Quantum Sensing Experiments


93. Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation


94. DS@GT ARC at CheckThat! 2026: LLM-Based Trace Ranking and Grouped Reward Modeling for Multilingual Numerical Claim Verification


95. Grounded in Consensus, In Step With Emerging Science: A Consensus-Anchored Multi-Corpus Clinical Chatbot for Long COVID


96. Authoring Agent Skills: A Software-Engineering Approach


97. CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models


98. Preliminary Guidelines for Using and Evaluating GenAI Tools to Support Systematic Literature Reviews


99. Stable FP4 Training via Transposition-Invariant Block Quantization


100. Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed


101. Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study


102. LLM as Forecasting Planner: Training-Free Text Conditioning for Time-Series Foundation Models


103. Beyond “What to Retrieve”: Uncertainty in Retrieval-Augmented Code Generation


104. Tokens are All You Need: Dual-purpose Semantic IDs for Achieving LLM-Level I/O Efficiency in recommendation systems


105. DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization


106. REPREC: Representation Driven Parameter-Efficient Recommendation System


107. MedJudgeRAG: Option-Wise Evidence Judgment with Dynamic Knowledge Graphs for Medical MCQA


108. Retrieval-Augmented Generation in LLMs for Mental Health: Quantifying the Incremental Contribution of Retrieval Within a Layered Safety Architecture


109. Multimodal Hybrid Retrieval-Augmented Generation for Scientific Document Understanding using Open-Source SLMs


110. Reading Without a Reader: Large Language Models Collapse Reading and Writing into a Single Entangled Code


111. From Naive RAG to Deep Agentic Retrieval: An Evolving Context Engineering Pipeline for Regulatory Compliance


112. Three Sides of Retrieval: Factorial Evidence for Document-Side, Query-Side, and Answer-Side Complementarity in RAG


113. The Effect of Text Chunk Size on Retrieval-Augmented Generation Performance


114. Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement


115. CARE-MH: Towards Unified, Reproducible, and Comparable Evaluation of Mental Health LLMs


116. DocAnnot – Accelerating the Creation of Key Information Extraction Datasets with GenAI-Powered Auto-annotation


117. Domain-Prior-Regularized Graph Modeling for Anomaly Detection in Cyber-Physical Systems