LLM 관련 주요 논문 - 2026-07-24

1. Beyond Sycophancy: Structured Resistance and Compliance in LLM Moral Reasoning


2. MIRROR: Learning from the Other View for Multi-Modal Reasoning


3. Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation


4. Agentic coding without the cloud: evaluating open-weight large language models on longitudinal data preparation tasks


5. Detecting LLM-Generated Tokens in Human–LLM Coauthored Text


6. PATS: Policy-Aware Training Scaffolding for Agentic Reinforcement Learning


7. Euclid-MCP: A Model Context Protocol Server for Deterministic Logical Reasoning via Prolog


8. An LLM-Driven Workflow for Automated Process Control Strategy Generation and Tuning from Dynamic Process Models


9. Safeguards for Speech2Speech LLM-Assistants: A Case Study in Automotive Applications


10. V-DEAL: Diagnosing Video Safety De-Calibration as an Understanding-Refusal Coupling Failure


11. AttriMem: Attribution-Guided Process Feedback for Agent Memory Learning


12. Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs


13. HiMe: Real-Time Self-Hosted Personal Agent Platform for Health Insights with Wearable Devices


14. EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization


15. Reexamining zero-shot summarization: Empirical investigation of trustworthiness of LLM-summarizers


16. GuardianAgentBench: Where Agents Fail and How to Guard Them


17. Beyond Independent Optimization: Compression, MoE Routing, and Quantization Interactions in Multimodal Edge Intelligence


18. Traceable Scholarship: Page Anchors and Ariadne’s Thread for Humanistic Inquiry in the Age of Generative AI


19. Is Deep Research Reliable? Misleading Knowledge Induces False Conclusions


20. Code Monitor Red Teaming for Public-Test-Passing Code


21. Auditing Evidence Use in Medical LLM Diagnosis


22. Auditing Provenance Sensitivity in LLM Agent Action Selection


23. Enhancing Explainable Cardiac Diagnosis with Guide-Grounded Multimodal LLMs


24. Profiling Lightweight Large Language Models


25. Refusal-Gated Decoding: Preserving Refusal Behavior Under High-Temperature Sampling


26. NVIDIA-labs OO Agents: Native Python Object-Oriented Agents


27. WaveformQA: Benchmarking LLM Temporal Reasoning on Digital Waveforms


28. KeySI: An Interaction Framework for Tuning Text Embeddings Based on Human Feedback


29. CMI-Mem: Toward Generalizable Long-Term Memory Management via CMI-Augmented Reinforcement Learning


30. AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use


31. DynamicMCPBench: A Trace-Grounded, Effect-Scored Benchmark for LLM Agents over Live MCP Servers


32. PromptPack: Scaling LLM Annotation Agents for Online Recommendation


33. Evaluating and Guarding Citation Faithfulness in Agentic Scientific Synthesis


34. ConfidenceBench: Evaluating Confidence Calibration in Large Language Models


35. Attention Degradation, Function Token Anchoring, and the Limits of Attention-Based Intervention in Large Language Models


36. Representation Robustness Under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving


37. Reliability-Aware LLM Alignment from Inconsistent Human Feedback


38. SiGMA: Sign-Guided Merging and Adaptation for Multimodal Continual Instruction Tuning


39. Telco-GAIA: Bilingual Benchmark for Agents in Telecom Domain


40. MiniCache: Reusable Program Caching with Small Model Interfaces for Efficient LLM Inference


41. Optimizing Hypergraph-Based RAG: Toward Better Fact Extraction and Chunk Retrieval


42. LeanFlow: A Case Study in Workflow-Driven Lean Autoformalization


43. MKEvolve: A Modular Multi-Agent Framework for Kernel Code Generation


44. FlowEdit: Information-Theoretic Control of LLM Reasoning Flows for Ill-posed Problems Involving Conflicts


45. ExecuGraph: A Multi-Agent, Execution-Grounded Framework for Reliable Backend Code Synthesis with Large Language Models


46. AISE-Bench: A Full-Cycle Curated Benchmark for Information Seeking on Academic Knowledge Graphs


47. Isolating LLM Alignment from Regex: Zero Coverage and Metric-Dependent Divergence Under Adversarial Mutation


48. EvoSQL: Memory-Augmented Critic-Generator Co-Evolution for Text-to-SQL


49. Autonomous Topology Mutation: Safe Runtime Restructuring for Multi-Agent LLM Systems with Capability, State, and Shadow Invariants


50. Directional Hallucinations: Ideological Drift in News-Grounded LLM Question Answering


51. Expectation Alignment of Language Models for Real-World User Expectations


52. The Devil is in the Spectrum: Mitigating Representation Collapse in LLMs via Topologically Regularized Side-Path


53. Tractable Hierarchical Control of Autoregressive Language Models


54. PersonaTrail: Benchmarking Personalized Web Agents through Browsing Trails


55. Routing Without Training: Controllable-Ratio LLM Offloading via Reliability Gating


56. Beyond Liars’ Bench: The Impact of Lie Typology, Depth, and Sparsity on Deception Detection in LLMs


57. Semi-Supervised Text-Attributed Graph Distillation


58. Benchmarking Large Language Models on Multi-Sensor Physical Hazard Assessment


59. SonicSampler: Unified Tile-Aware Kernels for LLM Sampling and Speculative Verification


60. VeriSimpl: Robust Optimization Modeling from Natural Language using Simplification-based Verification


61. Incomplete Prompt Jailbreaks in Large Language Models


62. Robust Critics: Defending LLMs Against Multi-Turn Attacks


63. Benchmarking the Personalization Capabilities of Large Language Models


64. PlanE: Meta Planning of Data, Tuning, and Inference for Extractive-based LLMs


65. DecodeShare: Tracing the Shared Subspace of LLM Decode-Time Decisions


66. InferenceBench: A Benchmark for Open-Ended LLM Inference Optimization by AI Agents


67. DC-Leap: Training-Free Acceleration of dLLMs via Draft-Guided Contiguous Leaping Decoding


68. Stochastic Sampling is Epistemically Shallow: The Dimensionality Gap Between Temperature Variation and Model Diversity in LLMs


69. ClickGuard: Detecting and Spoiling Clickbait News with Informativeness Measures and Large Language Models


70. Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts


71. AINTMA: Agentic AI Architecture for Autonomous Test Management with Generative Intelligence, Secure Cloud Communication and Adaptive Quality Analytics


72. 3D-Aware VLMs with Implicit and Explicit Geometries


73. From Resource Flow to Executable Tests: Petri-Net-Guided LLM Test Generation for Concurrent Stateful Rust APIs


74. Improved lower bounds for the Shannon capacity of odd cycles


75. Artificial Epanorthosis: Why large language models overuse a classical rhetorical figure, and how to mitigate it


76. Thinkink: 2D Spatial Ink-native Interaction with LLMs


77. When Are Reasoning-Based Guardrails Not Efficient? ResponseGuard: A Fast Vision-Language Guard for Real-Time Moderation


78. From Static Bibliometrics to Dynamic Knowledge Graphs: An LLM-Powered Framework for Modernizing Science, Technology, and Innovation (STI) Analytics


79. GRADRAG: Cross-Component Prompt Adaptation for Coordinated Multi-Agent RAG


80. Scaling Up Formal Representation of Clinical Trial Protocols in Ensemble Logic Using LLMs: A Preliminary Study


81. AI Assistants Overassist


82. Unlearning Under Imbalance: Benchmarking Fairness in Multimodal LLM Unlearning


83. A Comparative Evaluation of Embeddings and LLMs in a Greek Book Publisher Setting - The CUP Dataset


84. pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development


85. Case study: solving P-99 with LPTP and an LLM


86. Case study: proving sqrt(2) irrational with LPTP and an LLM


87. CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA


88. One More Turn, Less Regret: A Regret-Based Multi-Turn Benchmark for LLMs’ Clarification Policies


89. Training Large Language Models for Self-Explanation Faithfulness


90. Sparse Concept Channels in Frozen 3D CT Vision Encoders


91. Scientific exploration, collaboration and labor division in the large language model era


92. Beyond Heavy Log Curation: Perplexity-Based APT Detection via Unsupervised, Context-Augmented Language Models


93. The Geometry of Personality: Activation Steering with Jungian Cognitive Functions


94. Robostral Navigate


95. HARP: The Human–AI Research Platform


96. Are Diversity Metrics Measuring Diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles


97. IssueTrojanBench: Benchmarking AI Coding Agents Against Malicious Issue Requests


98. GPE: Evaluating Robust Evidence Aggregation for Fact Verification under Controllable GEO-Style Poisoning


99. Scaling Interpretable Transformers with Parity Bottleneck Layers


100. Demonstrating GenDB: Instance-Optimized and Customized Query Processing Code Generation via LLM Agents


101. Foundation-model-guided radiogenomic discovery linking cancer genomes to cancer scans


102. Geometric Configurations of Perturbed Jailbreak Prompts


103. StabilityBench: Benchmarking Instability in LLMs


104. SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales


105. ReliableTableQA:How Much Supervision Does Reliability Annotation Need?


106. Uncertainty-Aware Trust Estimation for Multi-LLM Systems via Structured Expert Judgement


107. PhantomFill: When the Form Demands an Answer, Language Models Invent One


108. Instruct-FD: Can Your Full-Duplex Speech System Follow Turn-Taking Instructions?


109. CAMeR: Keyword-Gated Hybrid Activation for Adaptive Memory Retention in LLM Agents


110. Dropping the Anchor: Statistical Context Summarization for Distributed Systems via Pulsar Attention


111. Learn2Zinc: Fine-tuning Small Language Models for Text-to-Model Translation in MiniZinc


112. RE-AD: Real-Time Requirement Adherence for Data Labeling


113. Response drift across frontier large language models


114. A Knowledge-Injection Framework for Zero-Shot Adaptation of LLMs to Delirium Prediction


115. The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs


116. Confidently Deceptive: How Confidence Amplifies the Risk of LLM Deception


117. Answer-then-Edit: Reasoning Skeleton Editing for Anti-Distillation with Preserved Utility


118. Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models


119. Making Open-Source Text LLM Watermarks Durable Against Merging


120. Break Through the Compression Bottleneck: From Theory to Practice


121. Moir: Let the Model Direct Its Own Story for Robust Cross-Domain Knowledge Editing


122. LLM-INSTRUCT at UZH Shared Task 2026: Constraint-Aware Retrieval and Selective Debate for Paragraph-Level Argument Mining


123. More Is Not More: What Matters for Diversity in LLM Opinions?


124. Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs’Hallucinations