LLM 관련 주요 논문 - 2026-06-04

1. AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety



3. Beyond Objective Equivalence: Constraint Injection for LLM-Based Optimization Modeling on Vehicle Routing Problems


4. BiasGRPO: Stabilizing Bias Mitigation in High-Variance Reward Landscapes via Group-Relative Policy Optimization


5. Inference-Time Vulnerability Beyond Shallow Safety: Alignment Along Generation Trajectories


6. FALSIFYBENCH: Evaluating Inductive Reasoning in LLMs with Rule Discovery Games


7. A Normative Intermediate Representation for ASP-Based Compliance Reasoning



9. Plan First, Judge Later, Run Better: A DMAIC-Inspired Agentic System for Industrial Anomaly Detection


10. Scaling Self-Evolving Agents via Parametric Memory


11. Simulate, Reason, Decide: Scientific Reasoning with LLMs for Simulation-Driven Decision Making


12. AgentJet: A Flexible Swarm Training Framework for Agentic Reinforcement Learning


13. Trivium: Temporal Regret as a First-Class Objective for Causal-Memory Controllers


14. Exploring Cross-Scenario Generality of Agentic Memory Systems: Diagnostics and a Strong Baseline


15. The Saturation Trap and the Subjectivity of Intervention Timing: Why Affect-Based Triggers and LLM Judges Fail to Time Interventions on Autonomous Agents


16. StepPRM-RTL: Stepwise Process-Reward Guided LLM Fine-Tuning for Enhanced RTL Synthesis


17. VAMPS: Visual-Assisted Mathematical Problem Solving Benchmark


18. Consensus is Strategically Insufficient: Reasoning-Trace Disagreement as a Knowledge-Representation Signal


19. SMAC-Talk: A Natural Language Extension of the StarCraft Multi-Agent Challenge for Large Language Models


20. Thinking Through Signs: PEEL as a Semiotic Scaffolding for Epistemically Accountable AI-Enabled Research


21. Toward Pre-Deployment Assurance for Enterprise AI Agents: Ontology-Grounded Simulation and Trust Certification


22. Failed Reasoning Traces Tell You What Is Fixable (But Not by Reading Them)


23. Towards Efficient and Evidence-grounded Mobility Prediction with LLM-Driven Agent


24. Audio Interaction Model


25. Arithmetic Pedagogy for Language Models


26. Automatic Generation of Titles for Research Papers Using Language Models


27. UniCAD: A Unified Benchmark and Universal Model for Multi-Modal Multi-Task CAD


28. Self-Reflective APIs: Structure Beats Verbosity for AI Agent Recovery


29. Invariant Gradient Alignment for Robust Reasoning Distillation


30. DAR: Deontic Reasoning with Agentic Harnesses


31. SharedRequest: Privacy-Preserving Model-Agnostic Inference for Large Language Models


32. From Agent Traces to Trust: Evidence Tracing and Execution Provenance in LLM Agents


33. Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning


34. Geometry-Aware Distillation for Prompt Tuning Biomedical Vision-Language Models


35. Provably Auditable and Safe LLM Agents from Human-Authored Ontologies


36. Learning While Acting: A Skill-Enhanced Test-Time Co-Evolution Framework for Online Lifelong Learning Agents


37. Description-Code Inconsistency in Real-world MCP Servers: Measurement, Detection, and Security Implications


38. Revisiting Vul-RAG: Reproducibility and Replicability of RAG-based Vulnerability Detection with Open-Weight Models


39. QO-Bench: Diagnosing Query-Operator-Preserving Retrieval over Typed Event Tuples


40. QuBLAST: A Framework for Quantizing Large Language Models with Block-Level Compression Approach and Activation Scaling Strategy


41. Ekka: Automated Diagnosis of Silent Errors in LLM Inference


42. Synthetic Personalities: How Well Can LLMs Mimic Individual Respondents Using Socio-Economic Microdata?


43. Multi-SPIN: Multi-Access Speculative Inference for Cooperative Token Generation at the Edge


44. Rollout-Level Advantage-Prioritized Experience Replay for GRPO


45. Temporal Order Matters for Agentic Memory: Segment Trees for Long-Horizon Agents


46. Dynamic Infilling Anchors for Format-Constrained Generation in Diffusion Large Language Models


47. ANN Search: Recall What Matters


48. GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling


49. Self-Evolving Deep Research via Joint Generation and Evaluation


50. Token Rankings are Unforgeable Language Model Signatures


51. MemoryDocDataSet: A Benchmark for Joint Conversational Memory and Long Document Reasoning


52. LoopMoE: Unifying Iterative Computation with Mixture-of-Experts for Language Modeling


53. Rethinking Sales Lead Scoring with LLM-based Hierarchical Preference Ranking


54. From Symbolic to Geometric: Enabling Spatial Reasoning in Large Language Models


55. MorphoQuant: Modality-Aware Quantization for Omni-modal Large Language Models


56. From Untrusted Input to Trusted Memory: A Systematic Study of Memory Poisoning Attacks in LLM Agents


57. Sparse Mixture-of-Experts Reward Models Learn Interpretable and Specialized Experts for Personalized Preference Modeling


58. Can I Take Another Dose? Evaluating LLM Decision-Making Under Temporal Uncertainty in OTC Dosing QA


59. Overview of the EReL@MIR 2025 Multimodal Document Retrieval Challenge (Track 1)


60. Recover-LoRA for Aggressive Quantization: Reclaiming Accuracy in 2-Bit Language Models via Low-Rank Adaptation with Knowledge Distillation on Synthetic Data


61. Supportive Token Revealing for Fast Diffusion Language Model Decoding


62. MM-BizRAG: Rethinking Multimodal Retrieval-Augmented Generation for General Purpose Enterprise Q&A


63. PerceptTwin: Semantic Scene Reconstruction for Iterative LLM Planning and Verification


64. A Systematic Analysis of Linguistic Features in AI-Generated Text Detection Across Domains and Models


65. EvalStop: Using World Feedback to Detect and Correct Reward Overoptimization in Multi-Tenant RLHF Platforms


66. Caught in the Act(ivation): Toward Pre-Output and Multi-Turn Detection of Credential Exfiltration by LLM Agents


67. Semantic Constraint Synthesis for Adaptive Trajectory Optimization via Large Language Models


68. SaliMory: Orchestrating Cognitive Memory for Conversational Agents


69. dMX: Differentiable Mixed-Precision Assignment for Low-Precision Floating-Point Formats


70. POLARIS: Guiding Small Models to Write Long Stories


71. Large Language Models Hack Rewards, and Society


72. Need to Know: Contextual-Integrity-Grounded Query Rewriting for Privacy-Conscious LLM Delegation


73. LLM Compression with Jointly Optimizing Architectural and Quantization choices


74. Spectral Scaling Laws of Muon


75. The Invisible Lottery: How Subtle Cues Steer Algorithm Choice in LLM Code Generation


76. RUBAS: Rubric-Based Reinforcement Learning for Agent Safety


77. LiftQuant: Continuous Bit-Width LLM via Dimensional Lifting and Projection


78. Unlocking Feature Learning in Gated Delta Networks at Scale


79. MaskForge: Structure-Aware Adaptive Attacks for Jailbreaking Diffusion Large Language Models


80. The Biomimetic Architecture of Software 4.0


81. CodegenBench: Can LLMs Write Efficient Code Across Architectures?