LLM 관련 주요 논문 - 2026-07-02

1. AutoMem: Automated Learning of Memory as a Cognitive Skill


2. Theoria: Rewrite-Acceptability Verification over Informal Reasoning States


3. Can Agents Generalize to the Open World? Unveiling the Fragility of Static Training in Tool Use


4. Agentic generation of verifiable rules for deterministic, self-expanding reaction classification


5. Graph-Native Reinforcement Learning Enables Traceable Scientific Hypothesis Generation through Conceptual Recombination


6. Self-GC: Self-Governing Context for Long-Horizon LLM Agents


7. AGI Maze as a Benchmark Framework for World-Modeling Agents


8. Agri-SAGE: Simulation-Grounded Multi-Agent LLM for Context-Aware Agricultural Advisory Generation


9. PHREEQC-MCQ-200: A Diagnostic Benchmark for Tool-Augmented Scientific Simulator Agents


10. Managed Autonomy at Runtime: Gear-Based Safety and Governance for Single- and Multi-Agent Cyber-Physical Systems


11. From Signals to Structure: How Memory Architecture Drives Language Emergence in LLM Agents


12. RareDxR1: Autonomous Medical Reasoning for Rare Disease Diagnosis Beyond Human Annotation


13. Making Failure Safe: A Constrained, Verifiable Agent Framework for Open-Web Data Collection


14. Measuring the Gap Between Human and LLM Research Ideas


15. Distill to Detect: Exposing Stealth Biases in LLMs through Cartridge Distillation


16. Diffusion-GR2: Diffusion Generative Reasoning Re-ranker


17. Adversarial Pragmatics for AI Safety Evaluation: A Benchmark for Instruction Conflict, Embedded Commands, and Policy Ambiguity


18. Skills Are Not Islands: Measuring Dependency and Risk in Agent Skill Supply Chains


19. Autonomous Scientific Discovery via Iterative Meta-Reflection


20. Towards Developing a Multimodal Chat Assistant for University Stakeholders: RAG-based Approach


21. CausalMix: Data Mixture as Causal Inference for Language Model Training


22. LongVQUBench: Benchmarking Long-Term Video Quality Understanding of Vision-Language Models


23. MemSyco-Bench: Benchmarking Sycophancy in Agent Memory


24. Behavior-Adaptive Conversational Agents: Toward a Fluid Personality Framework


25. Reading Order Inference for Complex Document Layouts


26. Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads


27. SWE-Doctor: Guiding Software Engineering Agents with Runtime Diagnosis from Multi-Faceted Bug Reproduction Tests


28. SenseWalk: Agent-Based Semantic Trajectory Simulation Powered by Large Language Models in Zoned Environments


29. From Personas to Plot: Character-Grounded Multi-Agent Story Generation for Long-Form Narratives


30. Recovering Input Text from Hidden States: Study of Gradient-Based Inversion of Decoder-Only Language Models


31. Exploring the Semantic Gap in Agentic Data Systems: A Formative Study of Operationalization Failures in Analytical Workflows


32. LeVLJEPA: End-to-End Vision-Language Pretraining Without Negatives


33. Phantom References: Hallucinated Citations That Survive Peer Review at Top-Tier Conferences


34. LLM-Guided ODE Discovery and Parameter Inference from Small-Cohort Aggregate Data


35. Self-conditioned Flow Map Language Models via Fixed-point Flows


36. LLVM-Bench: Benchmarking and Advancing Large Language Models for LLVM Compiler Issue Resolution


37. Loss Smoothing for Stable Adaptation Under Distribution Shift


38. Auditing Forgetting in Limited Memory Language Models


39. Active-GRPO: Adaptive Imitation and Self-Improving Reasoning for Molecular Optimization


40. BaseRT: Best-in-Class LLM Inference on Apple Silicon via Native Metal


41. MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos


42. Beyond the Prompt: Jailbreaking Function-Calling LLMs via Simulated Moderation Traces


43. Real-Time Hard Negative Sampling via LLM-based Clustering for Large-Scale Two-Tower Retrieval


44. NeuroCogMap Reveals Cognitive Organization of Large Language Models


45. DiscoLoop: Looping Discrete Embeddings and Continuous Hidden States for Multi-hop Reasoning


46. Mapping the Evaluation Frontier: An Empirical Survey of the Bias-Reliability Tradeoff Across Eleven Evaluator-Agent Conditions


47. An LLM-Based Framework for Intent-Driven Network Topology Design


48. What’s Hidden Matters: Identifying Planning-Critical Occluded Agents using Vision-Language Models


49. Testing Frontier Large Language Models’ Physics Literacy in Parallel Physical Worlds


50. SEFORA: Student Essays with Feedback Corpus and LLM Feedback Evaluation Framework


51. Adaptive Perturbation Selection for Contrastive Audio Decoding


52. EgoSafetyBench: A Diagnostic Egocentric Video Benchmark for Evaluating Embodied VLMs as Runtime Safety Guards


53. SLIM-RL: Risk-Budgeted Random-Masking RL for Diffusion LLMs Without Trajectory Slicing


54. GRPO, Dr. GRPO, and DAPO Are Three Operations on One Number: The Group-Standard-Deviation Identity


55. Harnessing the Latent Space: From Steering Vectors to Model Calibrators for Control and Trust


56. SWE-Router: Routing in Multi-turn Agentic Software Engineering Tasks


57. AGE: Adaptive-masking for Graph Embedding in Graph Retrieval-Augmented Generation


58. Prompting GPT-5 on Scrum Certification Questions: An Empirical Accuracy Study


59. Comparing Large Language Models on Scrum Certification-Style Questions: Accuracy, Stability, and Error Patterns


60. ATM: CID-Brokered Pre-Write Admission for Multi-Agent Code Co-Synthesis


61. LLMs in the Real World: Evaluating “AI” in Emergency Contexts


62. Learning User-Aware Recall: Personalized Retrieval in Long-Term Conversational Memory


63. Libra: Training the Environment for Agentic Information Retrieval


64. Towards an automated AI-based framework for floor plan compliance checks for residential buildings


65. PRA-RAG: Provably Robust Aggregation in Retrieval-Augmented Generation against Retrieval Corruption


66. SkillSelect-Serve: Budget-Controllable and QoS-Aware Skill Service Recommendation and Composition for Small LLM Agents


67. Prompt Optimization for User Simulation in Conversational Recommender Systems: A Multi-Objective Framework


68. Controllable Narrative Rendering for Enhanced Assisted Writing


69. SchemaRAG: Dynamic Large Schema Reduction for LLM-driven Structured Information Extraction


70. BaRA: BFS-and-Reflection Web Data Collection Agent


71. Persona Without Substrate: Regime-Dependence and the LLM Individuation Problem


72. From “Strings” to “Things” for Personal Knowledge Graphs: Evaluating LLM Triple Extraction for Recommendation Systems


73. UltraFlux: Data-Model Co-Design for High-quality Native 4K Text-to-Image Generation across Diverse Aspect Ratios