LLM 관련 주요 논문 - 2026-07-17

1. Pretraining Data Can Be Poisoned through Computational Propaganda


2. SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration


3. teLLMe Why (Ain’t Nothing but a Jam): Exploratory Causal Analysis of Urban Driving Data


4. When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space


5. Self-Evolving Human-Centered Framework for Explainable Depression Symptom Annotation


6. Can We Trust Item Response Theory for AI Evaluation?


7. Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy


8. Explaining Process Control Optimisation Recommendations via GradientSHAP and Implicit Differentiation


9. Contextualized Early Detection of Online Firestorms: A Sequential LLM-Based Approach


10. Transcoders for Investigating Deception in Language Models


11. AI vs Human Expert Reasoning: Assessing Agreements in Building Typology Predictions based on Street View Imagery


12. Project Kaleidoscope: Contextual, Human-Aligned Evaluation for Real-World AI Applications


13. SmartRAG: Native Graph-Based RAG for Mobile Device


14. TopoAgent: A Self-Evolving Topological Agent for Multimodal Scientific Reasoning


15. MCPEvol-Bench: Benchmarking LLM Agent Performance Across Dynamic Evolutions of MCP Servers


16. SportD: Can VLMs Physically Strategize?


17. MathCoPilot: An Interactive System for Human-AI Symbiotic Paradigm of Mathematical Research


18. Multi-LLM Collaborative MRI Report Generation for Visual Instruction Tuning in Brain Oncology


19. Collaborative Spatial Learning with Multi-LLM Agents in Networked Social Experiments


20. Alipay-PIBench: A Realistic Payment Integration Benchmark for Coding Agents


21. Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation


22. Towards an Intention Abstraction Layer for Autonomous Industrial Systems


23. Are LLM-Generated GPU Kernels Production-Ready? A Trace-Driven Benchmark and Optimization Agent


24. WrAFT: a Modularized Automated Writing Evaluation System for Argumentative Essays


25. RetroAgent: Harnessing LLMs to Search Over Structured Memory for Agentic Retrosynthesis Planning


26. VLT: A Vision-Language-Time Series Multimodal Foundation Model for Industrial Intelligence


27. Contextualized Evaluation of Vision Language Models through Dynamic, Multi-turn Interactions


28. SAGA: Schema-Aware Grounding for Agentic Text-to-SPARQL Generation


29. Step-Level Preference Learning for Generative Agents in Social Simulations


30. Reward-Free Evolving Agents via Pairwise Validator


31. Instrument Effects in Language-Model Honesty Evaluation: An Auditable Single-System Demonstration


32. CatalogAgent: A Supervisor-mediated Self-Learning System Enabling Context Engineering for GenAI Models


33. Chat2Scenic: An Iterative RAG-Based Framework for Scenario Generation in Autonomous Driving


34. CIPHER: A Decoupled Exploration-Selection Framework for Test-Time Scaling of Data Science Agents


35. Traccia: An OpenTelemetry-Based Governance Platform for AI Systems


36. Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions


37. AI Agents Do Not Fail Alone:The Context Fails First


38. Automatic Hard Example Synthesis with Multi-Level Agentic Data Curation


39. How Artificial Intelligence LLM Engines Shape the Global Conflict Information Environment


40. RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination


41. ReasFlow: Assisting Reasoning-Centric Scientific Discovery in Applied Mathematics via a Knowledge-Based Multi-Agent System


42. When a Verified World Model Still Loses: Play-Adequacy vs Prediction-Accuracy in LLM-Synthesized Code World Models


43. MemoHarness: Agent Harnesses That Learn from Experience


44. Orchestrating Power Grid Studies with Multi-Agent AI and MCP Servers


45. Enhancing Small Language Models Reasoning through Knowledge Graph Grounding


46. ToolAnchor: Anchoring Counterfactual Context to Boost Agentic Tool-use Capability


47. Human AI Construction of Bayesian Networks for Operational Decision Support – A Virtual Survey Approach


48. Interpretable Language Model for Closed-Loop Type 1 Diabetes Control


49. HG-RAG: Hierarchy-Guided Retrieval-Augmented Generation for Structured Knowledge Graphs


50. In-Place Tokenizer Expansion for Pre-trained LLMs


51. Symbal: Detecting Systematic Misalignments in Model-Generated Captions


52. MM-IssueLoc: A Controlled Benchmark for Evaluating Visual Evidence in Multimodal Repository-Level Issue Localization


53. Mask-Aware Policy Gradients for Diffusion Language Models


54. Digital Pantheon: Simulating and Auditing Coalition Formation with LLM Agents


55. OmniaBench: Benchmarking General AI Agents Across Diverse Scenarios


56. Show Me How You Reason and I’ll Tell You Who You Are: Reasoning Graphs for Robust LLM Authorship Attribution


57. StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows


58. Innocuous-Seeming Data, Latent Ideology: Ideological Generalisation in Finetuned LLMs


59. Does generative AI supersede supervised XMLC? A Benchmark Study on Automated Subject Indexing with German Scientific Literature


60. Can LLMs Build a MaxSAT Solver from Papers? The CoreForge Experience


61. Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality


62. Dialogue Summarization with Emotion Dynamics Using Topic- and Participant-Centric Decomposition


63. Large Audio Language Models for Spoofing-Aware Speaker Verification


64. Harnessing LLMs for Reliable Academic Supervision: A Comparative Study


65. An Intelligent-Cloud Edge Multimodal Interaction System for Robots


66. LLM-Driven Approach to Modeling Tool Interoperability in Automotive Domain


67. MemPoison: Uncovering Persistent Memory Threats and Structural Blind Spots in LLM Agents


68. Knowing You at First Glance: Inferring Apparent Personality from Faces


69. Answer-Conditioned Chains of Thought Degrade Verifiable-Reasoning Distillation in Large Language Models


70. SafeRelBench: A Spatial-Relation-Aware Benchmark for Process-Level Safety in VLM-Driven Embodied Agents


71. Controlled Reformulation Testing for Logical Consistency in Large Language Models


72. VTM-Nav: Hierarchical Visual-Topological Memory for Cross-Episode Object-Goal Navigation


73. Non-vacuous Generalization Bounds for Reinforcement Learning with Verifiable Rewards


74. Beyond Generalist LLMs: Specialist Agentic Systems for Structured Code Workflow Execution


75. Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel


76. HABIB_TAZ at SemEval-2026 Task 11: Disentangling Formal Logic from Content via Synthetic Training and Multi-Objective Optimization


77. Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values


78. Copy-on-Write Scoring: Application-Specific Agent Evaluations


79. Assessing AI in Introductory Physics Problem Solving


80. ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs


81. MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning


82. Closed-Loop Knowledge Dynamics: An Operational Framework for Saturation and Escape


83. Structured Feedback Improves Repair in an LLM Agent Loop


84. Towards Reliable AI-Assisted Analog Design: Template-Constrained LLM Agents for SAR ADC Generation


85. Certified Domain Consistency for Multi-Domain Retrieval: Label-Free Per-Domain Contamination Control with Conformal Risk Guarantees


86. Breaking Refusal in the First Half: A Mechanistic Study of the Prefill Jailbreak


87. CoEvoT: Co-Evolving Chain-of-Thought Prompting for Graph-LLM Reasoning


88. T5-CSBoost: Adversarial Perturbation Resistant LLM Fingerprinting


89. Information-Theoretic Limits of Reliability and Scaling in Language Models


90. Introspection Fine-Tuning (IFT): Training Small LLMs to Introspect


91. Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation


92. Eta Given Delta: Defining LLM Tool Efficiency With Marginal Tool Utility


93. Polestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs


94. Token Time Continuous Diffusion for Language Modeling


95. Automatically Evolving Prompt Guidelines for Task-Specific Optimization


96. LBA: Textual Hard-Label Adversarial Attack under Low Query Budgets


97. Just Keep Prompting: Evaluating Repetitive Socratic Prompting in VLMs