LLM 관련 주요 논문 - 2026-06-15

1. Towards Direct Latent-Space Synthesis for Parallel Branches in LLM-Agent Workflows


2. Abstracting Cross-Domain Action Sequences into Interpretable Workflows


3. Dense Coordinate-List Fine-Tuning Induces a Controllable Interference Surface in Vision-Language Models


4. From Chatbot to Digital Colleague: The Paradigm Shift Toward Persistent Autonomous AI


5. When the Tool Decides: LLM Agents Defer Blindly to Graph Neural Network Tools, and Stronger Backbones Defer More


6. GitOfThoughts: Version-Controlled Reasoning and Agent Memory You Can Replay, Diff, and Merge


7. Communication Policy Evolution for Proactive LLM Agents


8. AFFORDANCE20Q: Evaluating Affordance Reasoning from Physical Properties


9. SkillAudit: Ground-Truth-Free Skill Evolution via Paired Trajectory Auditing


10. Closing the Reflection Gap: A Free Calibration Bonus for Agentic RL


11. When Should Agent Trust Be Conditional? Characterizing and Attacking Skill-Conditional Reputation in Agent Swarms


12. VeriGeo: Controllable Geometry Question Generation with Numerical and Analytical Verification


13. FactoryLLM: A Safe and Open-Source AI Playground for Evaluating LLMs in Smart Factories


14. Formalizing Numerical Analysis: An Agent Pipeline and Quality Audit Beyond Kernel Acceptance


15. Minim: Privacy-Aware Minimal View for Agents via Trusted Local Sanitization


16. Adversarial Concept Search: Predicting Compositional Errors From Feature Geometry


17. Sorries Are Not the Hard Part: An Expert-Review Case Study of a Semi-Autonomous Formalization


18. Capability Minimization as a Safety Primitive: Risk-Aware Causal Gating for Least-Privilege LLM Agents


19. MA-ProofBench: A Two-Tiered Evaluation of LLMs for Theorem Proving in Mathematical Analysis


20. TwinBI: An Agentic Digital Twin for Efficient Augmented Interactions with Business Intelligence Dashboards


21. Orchestra-o1: Omnimodal Agent Orchestration


22. UP-NRPA: User Portrait based Nested Rollout Policy Adaptation for Planning with Large Language Models in Goal-oriented Dialogue Systems


23. ClinHallu: A Benchmark for Diagnosing Stage-Wise Hallucinations in Medical MLLM Reasoning


24. When Good Verifiers Go Bad: Self-Improving VLMs Can Regress on New Tasks


25. AudioDER: A Deduplication-Enhanced Reasoning Dataset for Post-Training Large Audio-Language Models


26. When Errors Become Narratives: A Longitudinal Taxonomy of Silent Failures in a Production LLM Agent Runtime


27. CARE: Controlling LLM-Generated Policies through Auditable Review of Evidence in Scientific Experimentation


28. SIMMER: Benchmarking Latent Failures in LLM Executable Planning with a World Model


29. From Shield to Target: Denial-of-Service Attacks on LLM-Based Agent Guardrails


30. tap: A File-Based Protocol for Heterogeneous LLM Agent Collaboration


31. Discovery under Hypothesis Redundancy: A Geometric Theory of Discovery Bottlenecks


32. I’m Sorry Driver, I’m Afraid I Can’t Do That: Appraising the Safety of LLMs within Automotive Contexts


33. From Prompts to Responses: Dual-Sided Data Leakage and Defense in Split Large Language Models


34. MeEvo: Metacognitive Evolution Combined with Natural Evolution for Automatic Heuristic Design


35. OdysSim: Building Foundation Models for Human Behavior Simulation


36. Implicit Reasoning for Large Language Model-based Generative Recommendation


37. Spatio-Temporal Audio Language Modeling for Dynamic Sound Sources


38. A Two-Stage Statistical Framework for Evaluating Associative Interference in Large Language Models


39. Hidden in Plain Sight: Benchmarking Agent Safety Against Decomposition Attacks with DECOMPBENCH


40. STREAM: Multi-Tier LLM Inference Middleware with Dual-Channel HPC Token Streaming


41. SANA: What Matters for QA Agents over Massive Data Lakes?


42. Mirage Probes: How Vision Models Fake Visual Understanding


43. SuperThoughts: Reasoning Tokens in Superposition


44. SpheriCity: Designing Trustworthy Conversational AI for Sustainability Decision Support


45. When Plausible Is Not Realistic: Evaluating Human Mobility in LLM-Based Urban Simulation


46. Aligning Quantum Operators with Large Language Models


47. A Benchmark and Framework for Evaluating Next Action Predictions in Spreadsheets


48. SEVRA-BENCH: Social Engineering of Vulnerabilities in Review Agents


49. VHDLSuite: Unified Pipeline for LLM VHDL Generation with Data Synthesis and Evaluation


50. HierSVA: A Data Synthesis Pipeline, Dataset, and Benchmark for LLM-Driven Hierarchical Hardware Formal Verification


51. An Agentic Retrieval Framework for Autonomous Context-Aware Data Quality Assessment


52. The Coin Flip Judge? Reliability and Bias in LLM-as-a-Judge Evaluation


53. Cross-Dataset Bloom Question Classification: Supervised Models and Prompted LLMs


54. GAGPO: Generalized Advantage Grouped Policy Optimization