LLM 관련 주요 논문 - 2026-06-29

1. Tandem Reinforcement Learning with Verifiable Rewards


2. JD Oxygen AI Item Center (Oxygen AIIC) V1: An Industrial-Scale LLM/VLM-Centric Solution for Item Understanding, Management, and Applications


3. NormAct: A Benchmark for Hidden Social Norm Compliance in Embodied Planning


4. Grounded Iterative Language Planning: How Parameterized World Models Reduce Hallucination Propagation in LLM Agents


5. Towards Reliable and Robust LLM Planning: Symbolic Feedback-Driven Iterative Self-Refinement Framework


6. ToE: A Hierarchical and Explainable Claim Verification Framework with Dynamic Multi-source Evidence Retrieval and Aggregation


7. DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums


8. Odyssey: Constructing Verifiable Local Truth-Preserving Foundation Models


9. Internalizing the Future: A Unified Agentic Training Paradigm for World Model Planning


10. When Does Personality Composition Matter for Multi-Agent LLM Teams?


11. AI-Model Network: Concept, Current State and Future


12. Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction


13. LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior


14. Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models


15. From Tokens to States: LLMs as a Special Case of World Models and the Continuous Path Beyond


16. Single and Multi Truth Data Fusion using Large Language Models


17. ToolPrivacyBench: Benchmarking Purpose-Bound Privacy in Tool-Using LLM Agents


18. MultiHashFormer: Hash-based Generative Language Models


19. Can LLMs Judge Better Than They Generate? Evaluating Task Asymmetry, Mechanistic Interpretability and Transferability for In-Context QA


20. Dialogue to Detection: A Multimodal Hybrid NLP Pipeline for Insurance Fraud Detection


21. From Black-Box to Clinical Insight: A Multi-Stage Explainable Framework for Speech-Based Cognitive Impairment Detection


22. Two-Stage Fine-Tuning for Protein Sequence Generation with Targeted Amino-Acid Composition


23. Agentic AI-Powered Re-Identification: An Emerging, Scalable Threat to Mobility Microdata Privacy


24. Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs


25. A Study of Temporal Fusion Strategies for Named Entity Recognition in Historical Texts


26. SpatialUAV: Benchmarking Spatial Intelligence for Low-Altitude UAV Perception, Collaboration, and Motion


27. Position Bias Correction is Insufficient for One-Pass Attention Sorting


28. NLL-Guided Full-Attention Layer Selection for Training-Free Sliding-Window Adaptation


29. Output-Space Allocation Costs for Calibration-Guided LLM Compression: An Empirical Study


30. Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models?


31. End-to-End Dynamic Sparsity for Resource-Adaptive LLM Inference


32. KG2Cypher: Data-Centric Pipeline for Building Enterprise Text-to-Cypher Systems


33. Bifocal Diffusion Language Models: Asymmetric Bidirectional Context for Parallel Generation


34. Enhancing Numerical Prediction in LLMs via Smooth MMD Alignment


35. Low-Agreeableness Persona Conditioning for Safe LLM Fine-Tuning


36. Mitigating LLM-based p-Hacking by Preregistering for the Next LLM


37. CBD: API-Only LLM Black-Box Unlearning through Controlled Behavioral Divergence


38. From Signals to Transfer: A Factorised Study of Probe-Based Uncertainty Estimation in Large Language Models


39. HybridCodec: Modeling Discrete and Continuous Representations for Efficient Speech Language Models


40. Dismantling Pathological Shortcuts: A Causal Framework for Faithful LVLM Decoding


41. On the Inseparability of Instructions and Data in Shared-Embedding Sequence Models


42. Benchmarking Multi-Modal Graph-based Social Media Popularity Prediction


43. Large Language Model Teaches Visual Students: Cross-Modality Transfer of Fine-Grained Conceptual Knowledge


44. Speculative Refinement: A Hybrid Autoregressive Diffusion Decoding Strategy and Its Behavior Across Benchmarks


45. Supersede: Diagnosing and Training the Memory-Update Gap in LLM Agents


46. SidConArena: An Environment Evaluating Agents in Open-Ended,Positive-Sum Bargaining Game


47. Agentic Publication Protocol: An Attempt to Modernize Scientific Publication


48. CalBrief: A Pilot Diagnostic Benchmark for Evidence-Calibrated Scientific Briefing with Large Language Models


49. Position: The Term “Machine Unlearning” Is Overused in LLMs


50. DataStates-LLM: Scalable Checkpointing for Transformer Models Using Composable State Providers