LLM 관련 주요 논문 - 2026-09-11

1. Can Edge-Deployable Vision-Language Models Identify Species?


2. Explainability Assistant: A Conversational XAI Interface for Interpreting Energy Consumption Models


3. From Parameters to Answers: How LLMs Retrieve and Use Their Internal Knowledge


4. When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making


5. COBRA-Skills: Contextual Bandit-Guided Evolution for Agent Skill Optimization


6. Geospatial AI, Dataverse Metadata, and the Study of Place-Based Government


7. MAPLE: Memory-Augmented Planning with Language and Evolution


8. Making Alternative Data Work: Context-Augmented LLMs for Financial Forecasting


9. Enabling Knowledge Graph Understanding at Scale with the EXplore Your Graphs ENgine (EXYGEN)


10. Characterizing Job Power Elasticity for Power-Flexible AI Training


11. ActMap: Single-Pass Uncertainty Quantification from Generation-Time Activation Maps


12. From Document Silos to Process Intelligence: A Multi-Layer Knowledge Graph for CMC Process Development


13. RouteRepair: Instance-Level Failure Diagnosis and Targeted Repair in LLM-Based Automated Heuristic Design for Routing Optimization


14. LLMs as Post-hoc Auditors of Physiological Plausibility in Symbolic Regression: A Clinician-Evaluated Case Study


15. Beyond Confidence: Stability-Aware Test-Time Adaptation for LLM Reasoning


16. Magenta: Closing the Loop Between Mathematical Reasoning and Lean Verification


17. Routing by Reasoning Need: Trajectory-Aware Decoding Control for Diffusion Vision-Language Models


18. Memory Compression for High-Fanout Agent Sandboxes


19. Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model


20. NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs’ Capability in Paper Novelty Assessment


21. A Voice-Interactive Multi-Agent System for Smart Operating Rooms: Architecture Design and Key Technologies


22. An AI-Powered Culturally Aware Chatbot for Stress Detection and Wellness Support among Pakistani University Students Using NLP and Machine Learning


23. Agentic Share-of-Search: A Multi-Agent AI System for Competitive Decision-Making in LLM-Mediated E-Commerce


24. Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment


25. SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics


26. Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial Sentiment


27. Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation


28. MOSAIC: Query-Aware Exploration Policy Adaptation for GraphRAG


29. Fork Where the Model Changes Its Mind: Belief-Shift Branching for Tree-Structured Reinforcement Learning


30. Demystifying the Privacy-Utility Trade-off in LLM Interactions


31. Decoupling Readiness from Release for Tail-Aware Scheduling of Agentic LLM Workflows


32. Studying Without a Syllabus: Task-Agnostic Environment Preprocessing


33. Towards a Deterministic Math Solver for Clinical Language Models


34. Domain-Specific Hallucination Detection in Large Language Models


35. Biology-in-the-loop: Amortized Adaptive Hit Discovery in CRISPR Screens


36. RetroThinker: Enabling Retrospective Thinking in Speech LLMs


37. Beyond Word Error Rate: A Switch Aware Evaluation of ASR and Audio Language Models on English Yoruba Code-Switched Speech


38. Recognizing Is Not Reversing: A Controlled Inversion Test of Fact-Preserving News Framing


39. LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation


40. ORCH: Organizational Principles Enable Collective Intelligence in Embodied AI


41. Language-Augmented Semantic Priors for B-Spline Surface Fitting


42. Ecdysis: Efficient and Effective Training of Runtime Harnesses for LLM Agents


43. Cross-Lingual Clinical Annotation Projection as Constrained Text Generation: A Six-Language Study


44. SWRouter: Similarity-Contractive Window Routing for Multi-Turn Large Language Model Conversations


45. X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation


46. On the Impact of Anonymization on the Performance of Large Language Models


47. E-CONAN (Entailment, CONtradition And Neutral) Benchmarks: Arabic Textual Entailment and Natural Inference Datasets


48. Your Model Already Knows Don’t Teach It, Learn to Ask It: Soft Prompting for Few-Shot Adaptation of Vision-Language Models


49. AI Soccer Analyst: Stage-Aware and Verifiable Human-AI Collaboration for Soccer Data Analysis


50. (Whose defaults?) Is artificial intelligence reorienting archaeological methods?


51. Beyond Benchmarks: Using VLMs to Reveal Systematic Classification Failures Under Real World Conditions


52. How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding


53. BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure


54. New Evidence, Same Choice: Testing Physical Experiment Selection in Vision Language Models


55. DeFiFusion: Combining Transaction Events with Smart Contracts to Detect Price Manipulation Attacks


56. Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training


57. ReactHuman: A Physics-Grounded Benchmark for Human-Like Reactive Decision-Making in Embodied Multimodal LLMs


58. Does Linguistic Structure Enrichment Enhance Coherence Assessment? Not With Current Architectures


59. DriftNet: A Dual-Head Trajectory Transformer for Detecting and Localizing Prompt Injection in LLM Agents


60. Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble


61. No-Box Vulnerability Analysis: Description-only Detection of Indirect Prompt Injection Vulnerabilities in MCP Servers


62. Beyond Static Guarantees: Measuring the Static-Pass Dynamic-Fail Gap in Security-Sensitive and LLM-Generated Python Code


63. Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu


64. When Synthetic Data Hurts: On Catastrophic Forgetting in Skill Retrieval for LLM Agents


65. CARTS: Contextual Autoregressive Rank Transcoding Steganography for Full-Capacity Keyed Text Encoding


66. Architecting the Secure AI-SOC: A Neurosymbolic Framework for Pipeline Integrity and Threat Mitigation


67. Beyond Verified Answers: Solver-Informed Self-Distillation for Bootstrapping Operations Research Language Models