LLM 관련 주요 논문 - 2026-08-28

1. CorporateBench: Large-Scale Q&A Benchmarking with Temporal Knowledge Bases


2. Sophistication in GenAI Use: Field Evidence from a Large Firm


3. LLMs Can Design Near-Optimal OR Algorithms


4. BrailleBench: Investigating Multi-Criteria Braille Comprehension in Large Language Models


5. What Makes Good Agentic Data? An ACE Lens on Data Generation for LLM Agents


6. Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable


7. When Tool Outputs Become Commands: Separating Action Induction from Runtime Authorization in Tool-Augmented LLM Agents


8. GRAIN: Bridging Name and Narrative Shifts in Real-World Graph Reasoning through Invariance-Rewarded Agentic RL


9. TransMeme: A Multi-Agent Framework for Cross-Cultural Meme Transcreation


10. LAAF: A Layered Accountability Architecture Framework for LLM Applications


11. pro-team at LLMs4OL 2026 Tasks Flagship and Reuse: Retrieval-Augmented Generation and Vocabulary-Constrained Filtering for Ontology Learning


12. Omni-Interactive Universal Embedder


13. DSA: Evidence-Aware LLM-Agent Orchestration for Multi-Market Stock Research


14. From Atomic to Agentic: Towards Interpretable Evaluation of LLMs’ Agentic Mathematical Capabilities


15. Counterfactual Bias Testing for Application Tracking System


16. Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall–workload trade-offs and run-to-run consistency


17. C-Unseen: Weak Signal Detection in Dynamic Temporal Knowledge Graphs via LLM Reasoning


18. BekchiAI: Measuring, Observing, and Controlling LLM Agents in One Click


19. LiveSim: Simulating Environment-Shaped Users in Multi-Agent Live-Stream Ecosystems


20. SymbolLKG: Towards Verifiable Logical Reasoning via Logical Knowledge Graph and Symbolic Solvers


21. Decoupling Planning and Control for Instructable Agents


22. AI Control Scientist: LLM-driven Agentic System for Automated Control Design


23. Discovering Relationships in Data Lakes Using Large Language Models: An Industrial Case


24. AgentFold: Closed-Loop Agentic Search for Protein Folding Model Design


25. Graph-Guided Selective Unlearning for Language Models: Controlling Support Routes Beyond Forget Seeds


26. Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training


27. AgentJudgeBench: A Multi-Difficulty Benchmark for Evaluating LLM Judges on Agentic Tool-Calling


28. DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows


29. Don’t Overthink, Don’t Underthink: Toward Adaptive Reasoning in Agentic AI


30. FaithSieve: Fine-Grained Evaluation of Math Proofs with Faithful Formal Evidence


31. Approved Too Late: Verdict Staleness in LLM-Guarded Self-Adaptive Systems


32. Assessing mentalization in humans and large language models


33. SKILL.state: Scalable Long-Horizon Agent Skills


34. The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts


35. LLM Agents for Time-Series: A Survey


36. Benchmarking AI Agents for Hardware Design Automation via MCP Tool Calling


37. Agentic AI for operating scientific instruments for nanoscale characterization



39. Is Your Neighborhood Safe? Place-based Stigma in Large Language Models’ Urban Safety Judgments


40. Why did My Robot Just Change Personality? Prompting Guidelines for a Grounded Robot Persona in LLM-Based HRI


41. AI Revealed Preferences


42. Refusal Is Not Robustness: Auditing Confident Fabrication in Large Language Models on a Provably Uninformative Clinical Pain Speech Transcript


43. A Task-Centric Ontology and Deterministic Domain Rules as a Verifiable Core for AI-Assisted Chemistry Problem Solving


44. GROUND: Reducing Hallucinations in LLM-Based Enterprise Analytics Through Governed Semantic Definitions


45. EEG-to-Report: An Annotation and Feature-Text Framework for Training Language Models on Clinical EEG


46. Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models


47. LLMs for Academic Workflows: An Evaluation of Literature Reviews Generated with Short and Long Context Windows of LLMs


48. CIFQA: A Deterministic Tool-Grounded Multi-Agent LLM Framework for Financial Query Answering


49. PICasso: An AI-Enabled Design Framework for Autonomous Optimization of Silicon Photonic Devices


50. Large Models for Battery Prognostics and Health Management: A Review and Future Roadmap


51. Standalone LLM and a Pre-specified Agentic Pipeline for Explaining ICU Mortality Predictions: a Feasibility Study on the eICU Demo Dataset


52. SWE-Prime: Fewer Trajectories, Better Performance


53. From Static to Dynamic: Benchmarking Real-World Code Review with MCR-Bench


54. RedEvoAgent: Automatic Red-Teaming Agent with Experience-Driven Skill Evolution


55. Persona-Execution Separation: An Architecture Pattern for Evolving LLM Agents under Execution Audit


56. How Language Models Organize and Structure Moral Knowledge


57. Making Clinical Language Models Auditable: Concept-Guided Fine-Tuning for Robust Prediction


58. Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit


59. Compositional Online Learning for Semantic Data Processing Systems


60. STEP: State-Aware Task Estimation and Planning with Multi-Modal LLMs for Human-Robot Collaboration


61. PACE: A Unified Condense-and-Extract Paradigm for Fast VLM Inference


62. LLMs in Digital EDA: A perspective on shifting roles from Generation to Orchestration


63. Safety Does Not Compose: Non-Decaying Loop State for Autonomous LLM Agents


64. Performance Foundations of Parallel & Distributed Reasoning Language Models


65. FaulT-Bench: Towards Benchmarking Network Troubleshooting LLM Agents under Unreliable User Tickets


66. When Memory Takes Gradients: Collaborative Vector Memory for Agentic Recommender Systems


67. PLCBench: Can Autonomous LLM Agents Turn PLC Access into Sustained Physical Impact?


68. From Reasoning to Pixels: Grounded Medical Multimodal LLMs for VQA and Segmentation


69. MedFG-VQA: Low-Frequency Memory and Graph Attention for Lightweight Medical VQA


70. Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research


71. Rethinking Message Passing as Retrieval for Text-Attributed Graph Learning


72. AesCanvas: A Large-Scale Dataset and Benchmark for Aesthetic Critique and Contextual Suitability


73. FOCUS & RePAIR: Mitigating Text Degeneration via Token-Level Guidance for Pruned Large Language Models


74. Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference


75. J-Zero: Unified Challenger–Solver–Judge Co-Evolution from Zero Data


76. Zero-Shot Self-Orchestration with Ledger-Based Control for Improved LLM Coding Performance


77. Diff Mining: Logit Differences Reveal Finetuning Objectives


78. Co-Evolving Structured Knowledge and Reasoning in Language Models


79. Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives


80. How Unlikely Is “Unlikely”? Assessing Verbal Probability Perception Across Large Language Models


81. Modality Maturity Index: A benchmark for assessing multimodal capabilities of omni models


82. MemToC: Benchmarking Memory-Tool Conflict Resolution in Large Language Models


83. How Do LLM Agents Actually Get the Flag? Trace-Level Provenance for Agentic Offensive Security Evaluation


84. NeuronFuzz: Safety Neuron Guided Fuzzing for LLM Safety Evaluation


85. A Reranker for Orchestrating Heterogeneous Speech and Text Retrievers


86. When the Canonical Completion Is Wrong: Formalizing and Measuring the Jump in Large Language Models


87. Investigating the Influence of Prompt and Response Languages on LLM Content Generation


88. PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants


89. A Multi-Framework Comparison of Outline Stages in Long-Form Generation with LLMs


90. Lost in Compression: A Controlled Cross-Lingual Audit of Extractive Prompt Compressors


91. DRL: A Deterministic Relational Middleware Layer for Transaction-Safe Enterprise NL2SQL Under Schema-Graph Scaling


92. Hallucinations in LLMs: A Lifecycle-Based Survey of Causes, Detection, Mitigation, and Prevention


93. Improving LLM Interpretability with User-Centric Chain-of-Thought Reasoning


94. Using Poly-Encoders for Computationally Efficient Automated Creativity Assessment


95. From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents


96. Mutual Debiasing via Dual-Seed Comparison for Probabilistic Sampling in Large Language Models


97. Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation


98. VFA: Empowering Multilingual MLLMs via Vision-Free Adaptation


99. Evaluating AI Generated Summaries for Cancer Patients



101. Beyond Accuracy: A Qualitative Analysis of Vision-Language Models for Hate Speech Detection in Memes


102. Position Is All You Need: A Free Lunch Token Compression Strategy for MLLM-based Referring Expression Segmentation


103. Syntax vs. Semantics: How Transformers Learn Deep Dependencies


104. From SQL to Knowledge Graphs: An LLM-Driven Multi-Agent Approach with Data Schema Improvement


105. Exploring the Role of LLMs in HPC Programming: A Survey