LLM 관련 주요 논문 - 2026-08-10

1. SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent


2. Blast Radius


3. PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents


4. Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing


5. A Picture is Worth a Thousand Tokens: How Vision Language Models Cut AI Energy Costs While Improving Accuracy


6. ResidencyRL: Reinforcement Learning in Simulated Clinical Environments



8. People Are Not Just Their Countries. Disentangling Social Determinants of LLM Value Alignment Across Europe


9. An End-to-End Agent Auditing Engine


10. WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN


11. Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models


12. Authoring and Management of Transparent Research Integrity Assessments of Randomised Clinical Trial Publications Using LLM-assisted Tools and Provenance Knowledge Graphs


13. EMAS: Stabilizing Multi-Agent System Evolution through Evidence-Guided Revision


14. Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory


15. NiyamAI - An Intent-Bound AI Agent with Cryptographically Verifiable Guardrails using Zero-Knowledge Proofs


16. A MARL Centered Reference Architecture for Large Language Model Augmentation in Smart Manufacturing


17. ReQuant: Fixed-Grid Discrete Refinement for Post-Training Quantization


18. Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation


19. Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints


20. Blind to the Pivotal Vote: Aggregate Independence Metrics Miss Where Verification Actually Helps


21. Science Edge Evaluation: SEE the Missing Step Toward Real Scientific Discovery


22. Deal Me Maybe: The Role of Emotions in Multi-Agent Negotiation


23. Long-Horizon Agent Trajectory Attribution: A Unified Benchmark and Fine-Grained Annotation Framework


24. CEDAR: Agent-Orchestrated Tree Search for Goal-Directed Optimization of Complex Systems


25. Gated-BEPO: Confidence-Gated Bellman Credit Assignment for Large Language Model Agents


26. Evolving Parallel Algorithm Portfolios via Potential-Aware Instance Generation with LLMs


27. Capek 0.5: An Execution-Centric Vision-Language Model for Embodied Intelligence


28. Mind the Gap: A Dual Knowledge Graph Framework for Unified Multi-task User Intent Inference


29. IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents


30. MolBioKG: Grounding Out-of-Graph Molecules in Biomedical Knowledge Graphs via Multi-Resolution Structural Anchoring


31. AgentPatch: Coarse-to-Fine Weak-Task Repair for Merging Agentic Multimodal Large Language Models


32. A Multi-Agent Framework for Automated Coarse-Grained Molecular Dynamics of Polymers


33. Shape Your Feed: An LLM-based Agentic System for Conversational Recommendation


34. NxN E-valuation: Hypothesis Certification via a Conformal CRT Null


35. Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques


36. Divergent Response Modes in Frontier Language Models Under Steering Pressure


37. WebGrader: Training LLMs for Web Development with Self-Evolving Programmatic Grader


38. Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Prunin


39. Interpretable Unsupervised Community Detection with LLM-Symbolized Structured Processes


40. EntropyMoE: Entropy-Aware Sparse Expert Routing for Tokenizer-Free LLMs


41. CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity


42. Strategy-first synthesis planning for complex natural products


43. Taxonomy-Driven Analysis of Open-Source AI Risk Mitigation Tools


44. SABRE: Scalable and Automated Benchmarking of VLMs under Stress


45. Diffusion LLMs as Targets and Adversaries: Mechanistic Safety Exploits


46. PACE: Primitive-Aware Code Evolution for Automated Algorithm Design


47. LSEAD: A Privacy-Preserving LLM-Based Speech Analysis Framework for Early Alzheimer’s Disease Screening


48. Geo-Spatial Concept Probing of Large Language Models: Abstraction, Compositionality, and Grounding


49. Natural Language Processing Psychometrics


50. Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination


51. SCALE: Scientific Concept Aggregation via LLMs and Embeddings for Fine-Grained Taxonomy Extension


52. Reading Copom’s Tone: A Weighted LLM Framework for Hawkish-Dovish Sentiment, Forward Guidance, and Uncertainty


53. Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research Publications


54. Toward a Causal Data Management Ecosystem for Decision Making and Agentic AI


55. Measuring Concept Content in Text from LLM Activations: ESG Evidence from Concept Vectors and Linear Probes


56. PHOENIX: Fine-Tuned SLM-Powered Autonomous Satellite Lifetime Extension via Predictive Self-Healing and Multi-Agent AI Recovery


57. Human-Centered Explainable AI for TinyML Edge Devices: A Pareto-Based Selection Framework with LLM-Guided Design


58. RoRA: Role-Oriented Regional Allocation for Visual Token Pruning in MLLMs


59. Beyond Text Matching: Towards Reference-Free Evaluation for Human-Oriented Binary Reverse Engineering


60. An Agentic Hybrid Top-Down and Bottom-Up Approach to Knowledge Graph Generation


61. GPTKB 2.0: Browsing, Querying, and Auditing a Disambiguated LLM-Derived Knowledge Base


62. PHASE-Tree: Modeling Character-State Evolution in Long-Horizon Role-Playing Dialogue


63. Same physical state, different collective dynamics: state encodings select synchronization outcomes in language-model agents


64. Debias in Text, Believe Your Eyes: Text-Anchored Cross-Modal Transfer for Visual Counter-Commonsense Reasoning


65. Ask-E: An Environment for Calibrated Question Generation


66. Calibrating WEAT Against Anisotropy: ZCA Whitening as a Geometric Pre-Processing Step for Embedding Association Tests


67. Georeferencing Non-Gazetteered Place Names using Biological Specimen Records


68. FedVAR: Prototype-Aligned Federated Framework for Video Anomaly Recognition


69. Multi-Agent Forensic Reasoning for Generalizable Deepfake Video Detection


70. Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry


71. FutureBridge: Token Selection Beyond Local Preference in Collaborative Decoding


72. Coupling Planning with Episodic Memory in LLM Agents for Software Issue Resolution


73. LoRAScan: Detecting Backdoor Prompts in Low-Rank Adapters for Large Language Models via Down-Projection Activation Spikes


74. HLSmith: An Expert-Guided Agentic Framework for C/C++-to-HLS Translation


75. Genotypic Triggers: Exposing Pharmacogenomic Blind Spots via Host-Specific Backdoors in Generative Antimicrobial Peptide Models


76. Progressive Content Refinement with Decaying Reward Joint LinUCB


77. Multi-Level Modeling of Large Language Model Inference Latency and Energy via Hybrid Analytical–Machine-Learning Predictors


78. Online Monitoring and Corrective Steering of Programming Agents


79. Characterizing the Quality Profile of AI-Generated C++ in Production


80. MI-MIDI: Mechanistic Interpretability of Text-to-MIDI Generation Models via Probing, Lenses and Steering


81. Cryptanalytic Extraction of Isolated Bias-Free GLU Feed-Forward Blocks by Antipodal Separation


82. Beyond “AI Language”: The case for the idiolectal nature of LLM output


83. Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events


84. CyberForge: Verified Vulnerability Injection at Repository Level for Cybersecurity Agent Training


85. WorldMark: A Plug-and-Play World Knowledge Interface for Cross-Host Language Model Watermarking


86. Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models


87. Agentic Planning for Symbolic Execution


88. TEXAS: Task-Expert-Aware Supervision for Downstream Mixture-of-Experts LLM Adaptation


89. Mitigating Scoring Bias in LLM-as-a-Judge via Random Number Generation