LLM 관련 주요 논문 - 2026-09-21

1. A Lie Detector Test for Language Models: Reading Knowledge a Model Won’t Reveal


2. AutoViewMem: Self-Configuring Orthogonal Views for Conversational Long-Term Memory


3. EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise


4. LLM-Generated Feature Pools for Time Series Anomaly Detection


5. ECG Mirage: Revealing and Mitigating the Underutilisation of ECGs in Vision-Language Models for Clinical Prediction


6. Accelerating Dense LLMs via L0-regularized Mixture-of-Experts


7. One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction


8. The Communication Bottleneck: A Round-Trip Study of Tree-Structured Expression Serialization in Language Models


9. PolyBridgeBench: Benchmarking Multimodal LLMs for Physics-Grounded Bridge Design


10. LogicTrack: Auditing Reasoning Trajectories of Large Language Models with Formal Logic Solvers


11. GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation


12. Offline Multimodal Large Language Models for Decision Support in Air Operations


13. Efficient Benchmarking in Production: A Study of an Evolving LLM Agent


14. PlaceReasoner-Beta: Reasoning-Driven Macro Placement and Benchmarking


15. CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition


16. Information-Gain Rewards over Diversity-Pruned Tests: GT-Anchored Verifier Co-Training for Reliable Code Generation


17. SpecOpt: Contact-Diff Reasoning for Agentic Molecule Optimization Toward Binding Specificity


18. Can Agents Design Better Chips with a Higher Level Abstraction?


19. Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake


20. TinyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers


21. Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models


22. Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing


23. CaLR: Causal Latent Revision for Robust Diffusion Reasoning


24. Attention-Aware Routing: Coupling Routing and Attention in MoEs


25. RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models


26. Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw


27. DiaVLo: Diagnosing Behaviours of Vision-Language Models


28. Bayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents


29. NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities


30. When Should a Failing Robot Ask? Initiating Corrective Human-Robot Dialogue from Audited Sensor Evidence


31. Detecting Pretraining Data in Large Language Models from a Free-Energy Perspective


32. Benchmarking the Explanatory Quality of Open-Weight Vision-Language Models in Face Recognition


33. Do Personality-Tuned LLMs Make Better Social Agents?


34. CIPL: A Channel-Aware Framework for Recoverable Privacy Leakage in LLM Agents


35. Samsone: A Family of Open Small Audio Language Models for On-Device Inference


36. When Steering Fails in Latent Reasoning: A Latent-to-Language Transition Gap


37. SynthDemo-RL: Breaking the Zero-Reward Barrier in VLA Adaptation with LLM-Guided Synthetic Demonstrations


38. Chinese Competitive Debating Dataset and Benchmark


39. Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30


40. Micro-Collaborative Poisoning: A Distributed Attack on RAG Systems


41. GameLogicBench: Evaluating Coding Agents on Runtime Game Logic with Tick-Level State Assertions


42. VidOmni-Bench: A Benchmark for Fine-Grained Video Understanding via Spatio-Temporal Event Verification across Complexity and Duration


43. HE-Guardrail: A Homomorphic Guardrail Against Jailbreak Attacks for Encrypted Large Language Model Inference


44. AgentVidBench: A Multi-Hop Video Question Answering Benchmark for Evaluating MLLM Agents


45. From Memory to Behavior: A Behavior-Aware Role-Playing Framework for Social Media Influencers


46. CESBench: Benchmarking Large Language Models on Cryptographic Engineering Security for IoT Devices


47. Beyond Exact Match: Task-Aware GRPO for Cross-Domain PCBA Visual Question Answering


48. Verify, Don’t Trust: Agentic Model Development for Video Discovery Retrieval at Scale


49. KnowDemo: Knowledge-Guided Robot Demonstration Generation from Human Videos


50. FOCAL-VLA: Subtask-Guided Geometry Distillation and Implicit World Modeling for Vision-Language-Action Models


51. SWE-Proof: Can Language Models Resolve Real-World Issues with Machine-Checked Proofs?


52. The Stochastic Shift: A New Evaluation Paradigm for Text-to-SQL with AI Operators


53. Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models


54. Aligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support Generation


55. How Much of a Real Workload Can LLM-Generated GPU Kernels Actually Reach?


56. BI-Agent and BI-Bench: Towards Automating End-to-End Business Intelligence