LLM 관련 주요 논문 - 2026-08-31

1. Logos: An Agent Harness on a Cross-Process Bus


2. Training Communication-Efficient Mixture-of-Experts Language Models with Layer Re-Configuration


3. Learning to Use Tools: Reinforcement Learning for Tool-Integrated Mathematical Reasoning


4. Program Learning with Verifiable Rewards: Symbolic Backpropagation for Post-Training LLMs


5. VERA-8B: Evidence-Grounded Audit Risk Reasoning from SEC Filings


6. RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents


7. EvoUndo: Recoverability-Constrained Self-Evolution for LLM Agent Harnesses


8. GRACE:Gradient-guided Coreset Selection for LLM Unlearning


9. Propagating construction-time knowledge quality into medical question answering: A framework grounded in clinical guidelines


10. MAIL: Memory-driven, Adaptive, Incremental, and Literature-grounded Framework for Hypothesis Generation in Chemistry


11. Finding Where the Buck Stops: An Automated Failure Attribution-Based Reflection Framework for Multi-Agent Collaboration


12. Regime-Aware Portfolio Management via Retrieval-Augmented LLM-Guided Expert Switching


13. Beyond Task-Only Matching: Personalized Skill Routing with Counterfactual Evaluation


14. REINS: Refusal-Enhanced Inhibitory Steering with Sparse Autoencoder Features


15. Stay Within Your Bounds: Distance-Guided Decoding for Guaranteed Context-Free Grammar Compliance


16. Expert Knowledge & Machine Understanding: Bridging Reactome’s Ontology with LLM Semantic Embeddings


17. The Shape of Power: A Multilingual Framework for Social Power Reasoning in Dialogues


18. Speculative Probing: LLM Monitoring at Speculative-Decoding Cost


19. String: An Agentic OS Where Every App Is a Markdown File


20. Coverage, Not Credit: Failure-Credit Routing of Zeroth-Order Perturbation Budgets Does Not Improve On-Pool Sample Efficiency for LLM Agents


21. Automated Analysis Framework for Multilingual Climate-Health Literature Based on Multi-Agent Large Language Model


22. When Evidence Shapes Collaboration: Knowledge-Conditioned Topology Generation for Multi-Agent Systems


23. SABER: Stability-Aware Early Exit for LLM Reasoning via Adversarial Branch Probing


24. When Teacher Guidance Misleads: Reward-Aligned On-Policy Distillation


25. Cross-Session Decomposition Attacks: Scaling Risk and Intent-Aligned Retrieval Defense


26. From Documents to Reasoning: A Validated Synthetic Data Pipeline and Semantic-Aware Fine-Tuning for Financial Numerical Reasoning


27. AI Alignment through a Game-theoretic Lens: A Survey


28. Resource Constraints and Performance in Agentic AI Systems


29. HyQuant: Hybrid-Precision Quantization for LLM Attention


30. See, Hypothesize, Validate: Multimodal Agentic Framework for Discovering Governing PDEs


31. CoRe-MoE: Compact Reusable MoE for Continual Multimodal Instruction Tuning


32. SpikeOPD: Stable On-Policy Distillation for Autoregressive Spiking Language Models


33. From Uncertainty to Clinical Risk: Severity-Aware Conformal Planning for Interactive Medical Diagnosis


34. RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests


35. CURA: Certified Runtime Alarms for Computer-Use Agents


36. CEDAR: Automata as Verifiable Interfaces for Language-Guided Embodied Action


37. Credo: Reusable Declarative Primitives for Agentic Workflows


38. Why Didn’t It Check? Unsupported Final Claims and Their Repair in Two Tool-Equipped Language Models


39. LongGuard: Mechanistic Analysis and Training-Free Mitigation of Long-Context Failure in Safety Guardrails


40. Thinking Costs Tokens: When More Structure is Worth the Price


41. CareGraph: An Auditable Hybrid AI Framework for Evidence-Grounded Personalized Longitudinal Health Intelligence


42. LLM-Augmented Causal Discovery: Probabilistic Fusion of Edge Existence and Orientation


43. Rating the Raters: Rasch Measurement Theory for LLM Evaluation


44. Time Capsule of Testable Human Knowledge: 41 Years of Jeopardy! in a Single Free Local Model


45. Learning a Size-Weight Frontier for Synthetic-Augmented Inference


46. An Enclosed Mode Is a Gauge Choice: Topology Relative to Reach in Certified Code World Models


47. LLM-Based Agents for Software and Systems Security: Approaches, Applications, and Assessment


48. How Proper Scoring Rules Shape LLM Forecasting


49. NL2AGBench: Benchmarking LLM Auto-Formalization for AlphaGeometry


50. ARC-CT: Anatomy-Routed Contrastive Vision-Language Learning for 3D Chest CT


51. LongPIBench: A Long-Context Benchmark for Prompt Injection


52. When Linguistic and Internal Confidence Diverge in Large Language Models


53. Layered LLM Defenses as an Ensemble: Access Tiers, Inference Cost, and the Measured Failure Correlation Between Defense Layers


54. Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss


55. MaCoPlanner: LLM-Assisted Manual-Compiled Task Planning with Proactive Safety Verification for Robotic Industrial Panel Operation


56. Embedding Models for Stance-Aware Argument Retrieval


57. Spatial-Semantic Reasoning using Large Language Models for Efficient UAV Search Operations


58. Training-free Suction Grasp Detection for Deformed Aseptic Cartons Using Vision-Language Models and Geometric Surface Scoring


59. Text Restoration of Ancient Documents with Language Models


60. Gen-TAS: A Generative AI-Aided Hardware-Software Task Allocation Framework for FPGA-GPP Heterogeneous Systems


61. Nested Byte-Level Vocabularies Are Cheap to Deploy and Expensive to Share: A Pre-Registered Negative Result


62. VICT: Verifier-Instrumented Credit Tracing for Long-Horizon LLM Agent Reinforcement Learning


63. Dynamic Alignment Compensation for Hallucination Mitigation in Large Vision-Language Models


64. SimpCue: Cue-Based Prompting for Multilingual Text Simplification


65. Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code


66. Twin Worlds: Equivariance-Based Abstention for Evidence-Grounded Reasoning


67. When Can Conditional Flow Matching Replace Pointwise Negative Log-Likelihood?


68. A Method for Layer Bit-Width Allocation in LLM Quantization via Performance Maximization Under a Quality-Degradation Constraint


69. CAITLYN: Can LLM Agents Autonomously Synthesize Defenses against Emerging Injection Attacks?


70. Not to Break, but to Attest: Adversarial Probes for Privacy-Preserving LLM Verification


71. LandingAgent: A Reference-Annotated Dataset and Agentic Generation Framework for Landing Pages


72. OpenStamp: A Watermark for Open-Source Language Models


73. FedEHR-Agents: Federated Agentic Optimization for Automated EHR Modeling


74. FISGuard: Defending Against Membership Inference via Fixed Input Subspaces


75. ContextLeak: Exfiltrating LLM Agent Context via Malicious Tools


76. First Make It Playable, Then Make It Good: Staged Interaction Learning for Small Dialogue-Game Agents


77. Knowing Before Answering: Decoding Language Models for Reliable RAG


78. LitCurate: A Configuration-Driven AI-Assisted Framework for Scientific Database Construction with an Application to Lower-Mantle Equation-of-State Data


79. Trajectory-Level Speculative Decoding for Diffusion Language Models


80. DAMP: Decay-Aware Mixed-Precision Recurrent-State Quantization


81. Quantization-Triggered Backdoors in Language Models: Cross-Quantizer Transferability and the Validation–Deployment Gap


82. A Survey on Rubric-Guided Reinforcement Learning for Language Models


83. Select, Don’t Train: The Benefits of Modular Entity Disambiguation with LLM-Based Selection


84. PACE: Publisher-Adaptive Content Extraction via Agentic Automation


85. The Effect of Emotional Context on Large Language Models’ Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models


86. Sledgehammer or Scalpel? A Fine-grained Adaptive Framework for Implicit Hate Speech


87. SciReC: Diagnostic Evaluation of Multimodal, Multi-Turn Relational Reasoning with Adaptive Interaction