LLM 관련 주요 논문 - 2026-09-04

1. Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints


2. Rethinking On-Policy Distillation of Large Language Models II: One Training Example


3. A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms


4. From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research


5. Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable


6. Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM


7. IRWOZ 2.0: A Large Language Model-driven Dialogue Dataset for Industrial Robot Conversations


8. FLY-EVAL++: An Evidence-Driven Evaluation Protocol for Safety-Constrained Flight Prediction with Large Language Models


9. InSituMeasure: Probing Situated Measurement Grounding in Industrial Scenes with Multimodal Large Language Models


10. LLM4CKD: Large Language Models for Early Stage Chronic Kidney Disease Screening


11. FiMI Banking: A Sovereign Model for Indian Retail Banking


12. Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting


13. Value-Preserving Architectures for Agentic AI Systems


14. STAIR (STructure Aware Information Retriever): A novel dataset and LLM based retriever for document structure augmentation


15. Bioinfoysis Technical Report


16. Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations


17. Semantic Bayesian World Models


18. SVG-Score: Human-Aligned Evaluation of Text-to-SVG Generation


19. SimSkill: A Lifelong Learning AI Agent for Autonomous Mastery of Traffic Simulation


20. Proactive Service Agents: A Unified Decision Framework, Methods, and Evaluation


21. Analysis of Prompt Engineering for Drug Toxicity Prediction


22. KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents


23. HalluPeer: A Taxonomy-driven Benchmark for Detecting Hallucinations in Scientific Peer Reviews


24. GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis


25. Dalek: A Constructive Agent Machine


26. NeoRed: A Knowledge-Logic-Alignment Multimodal Large Language Model for Neonatal Respiratory Disease Diagnosis


27. CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning


28. GrowPage: On-Demand KV Budgeting for Efficient LLM Reasoning Serving


29. Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models


30. DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents


31. Dude: A Dual-Detection Multi-Agent System for Paper-Code Discrepancy Detection


32. Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation


33. A Prompt-Engineering Approach to Develop Scalable, Flexible, and Real-Time Hybrid Micro-Level Personalization in a General Purpose AI Teaching Assistant


34. Fresh Memory, Stale Plans: Dependency-Scoped Validation for Distributed LLM-Agent Memory


35. Speculative Macro Commit for Faster Tool-Using Agents


36. MasterControl Seventeen Every Time


37. Seeing Before Synthesizing: VLM-Guided Transition Event Discovery for Weakly-Supervised Dense Video Captioning


38. Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views


39. SWE-Gate: Passing Functional Tests Is Not Enough for Software Engineering Agents


40. SENTINEL-RL: Offloading Topological Reasoning from LLM Agents in the Security Operations Center


41. When Models Edit Too Much: On the Fidelity of Minimal Code Edits


42. Representational alignment yields generalizable safety in language models


43. Investigating the Ability of Large Language Models to Analyze Recipes for Diabetes


44. GraFT: A Training-Free Framework for Spatial Reasoning in Multimodal Large Language Models via 3D Scene Graphs


45. A Blind Trust, the Bloody Thrust: When Attacker-Controlled Hook Updates Steer AI Agent Harnesses towards Malicious Behaviors


46. Free Pause Tokens


47. LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes


48. IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks


49. Beyond BLEU: A Case for Redefining Sign Language Translation Benchmarks


50. Can LLMs Extract Architectural Design Decisions from Source Code Commits? - A Preliminary Exploratory Study


51. ToolDF: Tool-Integrated Reasoning for Mixed-Authenticity Audio Deepfake Detection


52. Remember and Reweight: Enhancing Multi-Agent Debate with Experience Memory and Confidence Estimation


53. FailBench: How Reliable are VLMs at Judging Robot Task Success?


54. When Users Don’t Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents


55. It’s the Problem, Not the Path: Budget and Difficulty Confounds in LLM Reasoning Trajectories


56. StrixAE: An Intelligent Agent for Audio Enhancement under Complex Distortion Coupling in Real-World Scenarios


57. TabScope: Question-Adaptive Scope Selection for Table Question Answering


58. SHELF: A Synthetic Harness for Multi-Task Bibliographic Benchmarking


59. Privacy-Preserving Topology-Guided Safety for LLM-Based Multi-Agent Systems via Federated Graph Learning


60. When Optimization Becomes Manipulation: Defending Generative Search against Malicious Generative Engine Optimization


61. The Geometry of Ignorance: LLMs Know When to Temper Bayesian Priors


62. Privacy-Preserving Heterogeneous Multi-LLM Federated Inference for Cognitive Diagnosis


63. Reflect-SQL: A Self-Reflection Based Framework for Text-to-SQL


64. Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation


65. Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions