LLM 관련 주요 논문 - 2026-06-17

1. The Stanford EDGAR Filings Dataset: Reconstructing U.S. Corporate and Financial Disclosures into Layout-Faithful and Token-Efficient Pretraining Data


2. Learning Cardiac Electrophysiology Digital Twins Through Agentic Discovery of Hybrid Structure


3. WEQA: Wearable hEalth Question Answering with Query-Adaptive Agentic Reasoning


4. Trust the Right Teacher: Quality-Aware Self-Distillation for GUI Grounding


5. IsabeLLM: Automated Theorem Proving Applied to Formally Verifying Consensus


6. A Unified Framework for Context-Aware and Relation-Aware Graph Retrieval-Augmented Generation


7. Agentic AI-based Framework for Mitigating Premature Diagnostic Handoff and Silent Hallucination in Healthcare Applications


8. PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience


9. ProvenanceGuard: Source-Aware Factuality Verification for MCP-Based LLM Agents


10. LLM Consumer Behavior Theory: Foundations of a Novel Research Field


11. Small Initialization Matters for Large Language Models


12. How Inference Compute Shapes Frontier LLM Evaluation


13. PreAct: Computer-Using Agents that Get Faster on Repeated Tasks


14. DiagFlowBench: Evaluating How Language Models Handle Off-Procedure Inputs in Grounded Diagnostic Dialogue


15. StepGuard: Guarding Web Navigation via Single-Step Calibration


16. DecoSearch: Complexity-Aware Routing and Plan-Level Repair for Text-to-SQL


17. Shattering the Autoregressive Curse: Dynamic Epistemic Entropy Orchestrated Erasable Reinforcement Learning for LLMs


18. LongWebBench: Evaluating Structural and Functional Webpage Generation in Long-Horizon Settings


19. EComAgentBench: Benchmarking Shopping Agents on Long-Horizon Tasks with Distributed Hidden Intent


20. Using Cognitive Models to Improve Language Model Simulation of Human Persuasion Games


21. Beyond Domains: Reusing Web Skills via Transferable Interaction Patterns


22. FinAcumen: Financial Multimodal Reasoning via Self-Evolving Experience Memory Harness


23. Brick-DICL: Dynamic In-Context Learning for Automated Brick Schema Classification


24. Closing the Feedback Loop: From Experience Extraction to Insight Governance in Verbal Reinforcement Learning


25. Surrogate Assisted Pedestrian Protection Design via a Foundation Model Orchestrated Workflow


26. DeepInsight: A Unified Evaluation Infrastructure Across the Physical AI Stack


27. SEAGym: An Evaluation Environment for Self-Evolving LLM Agents


28. LLM-as-Judge in Education: A Curriculum-Grounded Marking Pipeline


29. Can LLMs Be CEOs? Benchmarking Strategic Resource Reallocation with Multi-Role Agent Simulation


30. MapSatisfyBench: Benchmarking Satisfaction-Aware Map Agents through Behavior-Grounded Implicit Decision Factors


31. Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems


32. Distributed General-Purpose Agent Networks: Architecture, Key Mechanisms, and Prototypes


33. MemTrace: Probing What Final Accuracy Misses in Long-Term Memory


34. Quantifying Consistency in LLM Logical Reasoning via Structural Uncertainty


35. Nothing from Something: Can a Language Model Discover 0?



37. ReproRepo: Scaling Reproducibility Audits with GitHub Repository Issues


38. RubricsTree: Scalable and Evolving Open-Ended Evaluation of Personal Health Agents across Health Memory and Medical Skills


39. A Red-Team Study of Anthropic Fable 5 & Opus 4.8 Models


40. IUU+DB: Tracking Illegal, Unreported, and Unregulated Fishing, Seafood Fraud, and Labor Abuse through LLM-driven Information Extraction



42. Towards Understanding and Measuring COGNITIVE ATROPHY in LLM Behaviour


43. Structural Role Injection in Handlebars-Templated LLM Prompts: Triple-Brace Interpolation, Delimiter Family, and the Limits of HTML Auto-Escaping


44. Querying an astronomical database using large language models: the ALeRCE text-to-SQL system


45. When LLMs Analyze Scars: From Images to Clinically-Meaningful Features


46. Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond


47. A Neuro-Symbolic Approach to Strategy Synthesis for Strategic Logics


48. SoftMoE: Soft Differentiable Routing for Mixture-of-Experts in LLMs


49. PearlVLA: Progressive Embodied Action-Plan Refinement in Latent Space


50. Trustworthy Self-Composable Big-Data-as-a-Service: An LLM-Orchestrated Multi-Agent Framework for Automated Data Engineering, AutoML, MLOps Deployment, and Drift-Aware Lifecycle Optimization


51. AnchorKV: Safety-Aware KV Cache Compression via Soft Penalty with a Refusal Anchor


52. A Framework for Evaluating Agentic Skills at Scale


53. LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams


54. MIVE: A Minimalist Integer Vector Engine for Softmax LayerNorm and RMSNorm Acceleration


55. Vision-language models for chest radiography do not always need the image


56. SegTME-UNI2: A Foundation Model-Based Framework for Generalisable Multiclass Cell Segmentation and LLM-Driven Tumour Microenvironment Characterisation in Histopathology


57. See First, Answer Later: Visual Evidence Pre-Alignment via Sufficiency-Driven RL


58. FacProcessTwin: An LLM-Based System for Process Twin Development


59. TuneAhead: Predicting Fine-tuning Performance Before Full Training Begins


60. Divide, Deliberate, Decide: A Multi-Agent Framework for Fine-Grained Egocentric Action Recognition


61. Understanding LLMs in Title-Abstract Screening: From Disagreements to Recommendations


62. Visored: A Controlled-Natural-Language Prover for LLM-Generated Mathematics


63. LLM Features Can Hurt GNNs: Concatenation Interference on Homophilous Graph Benchmarks


64. Reinforcing Dual-Path Reasoning in Spatial Vision Language Models


65. OmniDrive: An LLM-Choreographed Multi-Agent World Model with Unified Latent Co-Compression for Multi-View Driving Video Generation


66. Scaling Enterprise Agent Routing: Degradation, Diagnosis, and Recovery


67. Unlocking LLM Code Correction with Iterative Feedback Loops


68. Online LLM Selection via Constrained Bandits with Time-Varying Demand


69. AIPatient Arena: EHR-grounded evaluation of large language models in end-to-end clinical consultation workflows


70. AUTOGATE: Automated Clock Gating via Toggling-Aware LLM-based RTL Rewriting


71. MODE-RAG: Manifold Outlier Diagnosis and Energy-based Retrieval-Augmented Generation Evaluation


72. Patients With Personality: Realistic Patient Simulation through Controlled Diversity and Selective Disclosure


73. Enhancing Pathological VLMs with Cross-scale Reasoning


74. SoK: AI-Augmented Binary Reversing


75. NarrativeWorldBench: A Frontier-Saturated Benchmark and a Latent World Model for Long-Horizon Co-Creative Audio Drama


76. Visuals Lie, Consistency Speaks: Disentangling Spatial Attention from Reliability in Vision-Language Models


77. Model Validation of Agentic AI Systems: A POMDP-Based Framework for Belief-State, Forecast, and Policy Validation


78. DriveJudge: Rethinking Autonomous Driving Evaluation with Vision-Language Models


79. Do Large Language Models Always Tell The Same Stories?


80. Rift: A Conflict Signature for Deception in Language Models


81. Trust-Aware Multi-Agent Traceability: Confidence-Calibrated Knowledge Graphs for Consistent Software Artifact Management


82. PowerOPD: Stabilizing On-Policy Distillation with Bounded Power Transformation


83. Cluster-Aware Dual-Level Test Specification Generation for Large-Scale Automotive Software Requirements


84. Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference


85. LineageMark: Multi-user White-box Watermarking for Contribution Tracing in Model Derivation Chains


86. MODE: Modality-Decomposed Expert-Level Mixed-Precision Quantization for MoE Multimodal LLMs


87. An Evaluation of Data Leakage Risks in Tool-Using LLM Agents in Realistic Scenarios


88. Prefill/Decode-Aware Evaluation of LLM Inference on Emerging AI Accelerators


89. ParkingTransformer: LLM-Enhanced End-to-End Trajectory Planning for Autonomous Parking


90. Extracting Semantics: LLM-Guided Automatic Population of Robot Ontology from URDF


91. Towards Distributed Inference of LLMs on a P2P Network


92. Correct When Paired, Wrong When Split: Decoupling and Editing Modality-Specific Neurons in MLLMs