LLM 관련 주요 논문 - 2026-09-23

1. Grow the Harness, Not the Context: From Strategy-Free Scaffolds to Reusable Specialist Agents


2. JEV-as-a-Judge: Accept When Confident, Escalate When Unsure


3. REFLEX with Jev for Efficient Selective Control in LLM Agents


4. Identifying Intelligent Processes via Online Sequential Testing


5. EADC: Evaluation of Advanced and Deep-level Compliance in Large Language Models


6. MAC-RRG: Iterative Multi-Agent Collaboration for X-ray Radiology Report Generation


7. DTOC: Dynamic Tool Output Compression for Adaptive Context Management in AI Agents


8. Selection-Invariant Communication Compilers for Privacy-Aware Multi-Agent LLM Workflows


9. ChainUQ: Reasoning Consistency-Aware Uncertainty Quantification for Large Language Models


10. CQ4OE: A benchmark for assessing LLM-assisted ontology generation from competency questions


11. AgenticSizing: A Large Language Model-based Multi-Agent Framework for Analog Circuit Sizing


12. Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts


13. The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks


14. The Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural Variance


15. LingLan: An Advancing Traditional Chinese Medicine Diagnosis LLM with Multimodal Data


16. TCMaster: Confidence-Aware Querying and Workload-Guided Physical Design for Multi-Source Traditional Chinese Medicine Knowledge Graphs


17. How Strongly Should Task State Influence an LLM Agent?


18. Toolcompass: Guiding Tool Trialing, Not Suppressing It


19. Testing-Driven Reliability Audit of Trajectory-Based Early Outcome Prediction for LLM Agents: Target-Specific Calibration Transfer Persists Within a Single Benchmark


20. Ladders of Thought: A Self-Evolving Curriculum of Progressively Simplified Reasoning Traces


21. ChatT2: An Adaptive Framework for Developing a Large Language Model-Based Agent for Natural Product Domain Research


22. Reasoning-Preserving Fine-Tuning of Post-RL LLMs with Null-Basis LoRA


23. ArticleMiner: Ontology-Guided Knowledge Graph Construction from Scientific Publications


24. Direct Optimization of Generators for Search in Automated Theorem Proving


25. A Behavioral Trait Leaks into Preferences: Diagnosing Trait Interference in LLM User Simulators


26. Recovering Agentic Sovereignty: Mitigating the Consensus Paradox via Contrastive Epistemic Decoding


27. Spectra: A Rules-Driven LLM Pipeline for Automated KYC Document Processing


28. RAG-NAROK: Retrieval-Aware Knowledge Corpus Poisoning in RAG with Source-specific Refutation


29. Rollout Efficiency in Reinforcement Learning for Reasoning Large Language Models: A Taxonomy and Future Directions


30. Clarification Is Not Correction: LLMs Fail to Let Go


31. Potential for Enhanced Learning in Machine Learning Classes by Using Wiki LLM Indexing


32. Attention as a Routing Graph: Live Circuit Extraction from a Single Forward Pass


33. When LLM Agents Fail to Read the Room: ReAdapt for Relational Social Reasoning


34. The AI Neuroscientist: An Interactive Agentic Interface for Neuroimaging Analysis


35. Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation


36. SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue


37. Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen


38. Does AI Save Time on Product Design? A Randomized Controlled Experiment of AI Prompt-to-Design Workflows


39. Beyond Repeated Sampling: Learning Search Policies for LLM Reasoning


40. From Alignment to Access Control: A Framework for GenAI Policy Enforcement


41. Capable yet Parsimonious: Extracting and Characterizing Hidden Chain-of-Thought in Frontier Models


42. Greedy Decoding Is Not Precision-Invariant: Cross-Precision Output Divergence in LLM Inference


43. Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models


44. FeatLens: Feature-Guided Dynamic Code Graph Construction and Retrieval for Repository-Level Code Generation


45. TimeInteract: Towards Real-Time Interactive Intelligence for Streaming Time Series


46. PACT: From Credit Assignment to Critic Alignment


47. TransBERT: A Framework for Synthetic Translation in Domain-Specific Language Modeling


48. CompKV: Compensation-Aware KV Selection for Long-Context LLM Inference


49. On the security and privacy of LLMs in Mobility


50. Reducing Hallucinations in Large Language Models Through Integrated Self-Verification and Retrieval-Augmented Generation


51. WatchPoint: Executable User Feedback for Real-World Agentic Web Development


52. Dynamic Deep Prompt Optimization for Defending Against Jailbreak Attacks on LLMs


53. Silent Sabotage: Internal State Triggered Backdoor Attacks on LLM-Powered Robotic Systems


54. Magnitude Profile Pruning: Calibration-Free Structured Attention Head Removal for Transformer Compression


55. The Uncontrolled Variable: Vision-Language Model Refusal Responds to Image Presence in Ways Risk Cannot Explain


56. StepTrigger: Contact-State-Triggered Backdoor Attacks on VLM-Powered Legged Robots


57. TSS: Target-Side Sparsification for Speculative Decoding in Domain-Specific Large Language Models


58. CricRAG: Retrieval Augmented Vision-Language Models for Personalized Cricket Coaching


59. REVE: Efficient Hallucination Correction for Large Audio-Language Models via Reused Encoder States


60. BAS-OPD: Budget-Aware Selective On-Policy Self-Distillation for Fine-Grained Multimodal Perception


61. Syndrome, Synergy, and Safety: Structured Reasoning and Knowledge-Driven Alignment for TCM Prescription Generation


62. Slow Decay and Silenced Expression: Iterated Subliminal Trait Transfer in Language-Model Lineages


63. From Experts to Sub-experts: Fine-grained Parameter-Efficient Fine-Tuning for MoE LLMs


64. Deflecting the Value Compass: Interacting with Large Language Models Temporarily Shifts Human Value Priorities Toward Personal Focus


65. Compressing Long Context into Answer-Aligned Memory Embeddings for LLM Inference


66. Hill Sampling for Test-Time Scaling: A Simple and Better Alternative to Repeated Sampling, Evolution, and Training


67. Universal Fractal Natural Language Decision Map: Real-Time Edge Triage Across Heterogeneous Domains


68. RGSQ: Riemannian Geometry-Sensitive Quantization for Large Vision-Language Models


69. Terminal Shrinkage Averaging Reveals a Schedule-Estimator Interaction in LLM Pretraining


70. Passes Alone, Fails Together: Benchmarking Semantic Coordination in Parallel LLM-Agent Development


71. Trains but Doesn’t Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers


72. Indirect tipping: a social attack surface in AI agent populations


73. From Pattern Recognizers to Personalized Companions: A Survey of Large Language Models in Mental Health


74. The Probabilistic Structure of Large Language Models


75. Rachel: A general-purpose language model directs and revises retrosynthetic routes


76. LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay


77. LLM-Driven Training-free Location-Attribute Synergic Fusion: A Closed-Loop Paradigm for Dual-source Encrypted POIs and LULC Mapping


78. Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibratione



80. “As a Language Model…”: Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It


81. Not All 4-bit Quantizers Are Equal: Deployment-Time Mitigation of PII Leakage in Fine-Tuned Small Language Models