LLM 관련 주요 논문 - 2026-08-12

1. V-FiLLM: Verified Financial LLM Reasoning Benchmark


2. IO Factory: Simulating AI-Enabled Influence Campaigns at Scale


3. Hypothesis Frontier: Verifier Guided LLM and Symbolic Search for First-Order Induction


4. EvoMem: Memory-Augmented Evolution for Code Optimization


5. SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation


6. Rule of Thumb: Explaining Artificial Intelligence Systems using Partial Information


7. Self-Correcting Long-Horizon Search Agents via Tree-Structured Memory


8. REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems


9. VERDICT: Training-Free Step-Wise Verification of Multimodal Reasoning via Disagreement-Aware Consensus


10. HexEval: An Evidence-Driven Hexagonal Framework for Multidimensional Scholar Assessment


11. SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models


12. Measuring Semantic Abstractness of SAE Features via Nonlocality


13. RadFusion: Towards Threshold-Controllable Radiology Report Generation


14. MEGA: Self-Evolving Agent Optimization Infrastructure via Wisdom Graph


15. From Faulty Memories to Corrected Actions: Dependency-Guided Rollback Repair for Memory-Augmented Agents


16. GeoForge: Non-Parametric Self-Evolving Agents for Earth-Observation Reasoning


17. INSIDE the Student’s Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators


18. Predicting Space Groups of Double Perovskites by LLM with Dynamic Few-Shot Learning


19. Multi-Granular Rationale-Guided Molecular LLM for Property Prediction


20. Evaluating Rational Contracting in Natural Language


21. RLMOpt: Adaptive Prompt Optimization via Recursive Language Models


22. Continuous Interaction Diffusion: A Diffusion-Native Runtime for Asynchronous Tool-Augmented Reasoning


23. Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance


24. Recovering Wasted Compute in Autoresearch Agents


25. Hierarchical Compositionality for An Assistive AI Agent


26. Toward a Theory of Value in AI Alignment


27. Logit-Boundary Geometric Belief Interfaces and Sparse Sheaf-Enclave Protocols: A Self-Contained Substrate for Secure Network Electronic Health Record (EHR) Interoperability


28. Interpreting Language Model Hidden States at Scale


29. Beyond Detection: Evaluating Defensive LLMs Against AI-Generated Social Engineering in Live Turn-by-Turn Interaction


30. Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems


31. Decodable But Not Detachable: Training Data Granularity Determines Parametric Modularity in Large Language Models


32. Evaluation-Conditioned Training: Teaching Models to Generalize to Stronger Oversight Regimes


33. Mitigating Bus Bunching with Reinforcement Learning Enhanced by Semantic Stop Embedding


34. Post-Hoc Sparse Coding of Latent Communication Between Vision-Language Model Agents


35. TRACE: Trustworthy Retrieval-Augmented Conversational Engine


36. Generating Attacks for LLMs with GFlowNets


37. CHORUS: Complementary Experts for High-Coverage Testbench Stimulus Generation


38. Closed-Loop LLM Co-Pilots for Digital Agriculture


39. ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls


40. Attention-Path Fragility as an Uncertainty Signal in Large Language Models


41. Quantum Coordination Advantages in AI State-Tracking Tasks: Semantic Compilation and Latent Memory


42. Workflow Cards: Structured Summaries of Workflow Executions Using Provenance Data


43. ReLTEx: Reliable LLM-based Taxonomy Expansion


44. CARE: Confidence-Aware Reasoning for Reliable Medical VQA


45. Evidence-Grounded Trustworthy Multimodal Reasoning and Evaluation Benchmark in Complex Urban Scenes


46. A Cost-Efficient Routing Pipeline for Multilingual Short-Text Classification Using Small Language Models


47. Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation


48. FaithformBench: Benchmarking Faithfulness of Mathematical Chain-of-Thought Autoformalisation


49. VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?


50. Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition


51. Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation


52. A Gateway Architecture for Enterprise MCP Authentication: Unifying Heterogeneous Auth, Identity Delegation, and the User / Non-User Persona Problem


53. Most biomedical publications show signs of LLM-assisted writing


54. Conversational Orchestration for Organic 6G


55. Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control


56. ProTAGAD: A Foundation Model for TAG Anomaly Detection with Decoupled Topological and Textual Prototypes


57. Optimize Cheap, Deploy Strong: Cost-Aware Cross-Tier Transfer for Evolutionary Optimization


58. SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information


59. MedUP: Awakening Unified Understanding and Perception in Medical Vision-Language Models


60. Compute-Optimal Is Not Cluster-Optimal: Systems-Aware Scaling for Sparse Mixture-of-Experts


61. DegradeQuery: Counterfactual Tuple Pretraining for Context-Aware PROTAC Degradation Prediction


62. ImpactHO: Importance-Aware KV Cache Transfer for Multi-User Edge LLM Handover


63. On Understanding, Identifying, and Mitigating Vulnerabilities in Agentic Large Language Models


64. Dynamic Context Adapters: Efficiently Infusing History into Vision-and-Language Models


65. Unlocking the Power of Medical Tabular Data via Semantic-Aware Multimodal Pre-training


66. SafeCap: Improving LVLM Safety with Image Captioning Reinforcement Learning


67. Lost in Reconstruction: Aligning Action Representations with Language in Vision-Language-Action Models


68. MD-ProTector: Positioning Multiple Data-Driven Prototypes for LLM-Generated Text Detection


69. Towards Efficient Reasoning in LLM-Based Recommender Systems via Model Merging


70. From Reasoning Depth to Reasoning Breadth: Evaluating Multi-Point Associative Reasoning in Large Language Models


71. Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique


72. Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models


73. Persona Conditioning as an Assessor-Sensitivity Probe for LLM-Based IR Evaluation


74. Beyond Forecasting: Recasting Volatility Control as a Routing Problem


75. MemSpec: Memory-Aware Runtime for Adaptive Draft Scheduling in Speculative Decoding on Edge Devices


76. Expert-Guided g-computation with Large Language Models for Estimating Causal Effects on Timings: Applications to Hospital Quality Improvement


77. Do Personalized Skills Help Coding Agents? An Empirical Study of Developer Interaction Histories


78. Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility


79. Comprendia: AI-Augmented Code Comprehension


80. Withholding the Completing Chunk: Deterministic Pair-Completion Guardrails for Streaming LLM Output


81. Locally Deployable Small Language Models for Emergency Department Decision Support: A Systematic Benchmark of Fine-Tuning Strategies


82. Toward Human Rights Benchmarking for LLMs: A Pilot Methodology


83. TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent


84. ELMER: Evolutionary Language Model that Explores and Refines


85. The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse


86. Multimodal Item Parameter Estimation using Simulated Response Probabilitie


87. Procedural Fairness Failures in RLHF from Preference Averaging


88. Finding the Signal in the Spam: Jointly Learning Rewards and Worker Reliability from Pairwise Comparisons


89. DOCSCHISEL: Adaptive Tool Documentation Optimization Framework for LLM Agents


90. Eleven Years of BRACIS: A Meta-Scientific Study of the Brazilian Conference on Intelligent Systems


91. HoosierHelp: Benchmarking LLM Agents for Social Service Navigation


92. When Chain-of-Thought Helps and When It Hurts: An Empirical Investigation of the Serial-Depth Bottleneck in LLM Reasoning


93. How to Dogfood Your AI Chat Agent: A Three-Layer Evaluation Framework with Goal-Directed NPC Simulation


94. LLM Agents Factory: Retrieval of Domain-Specific LLM Agents


95. “YES! YES! I absolutely love this insight!” Affirmative Narration as Interactional Strategy in Dialogues with LLM Chatbots