LLM 관련 주요 논문 - 2026-07-23

1. PoTRE: Test-Time Reasoning inspired by Cognitive Heterogeneity


2. CUSUM-Shaped Inference-Time Monitoring and Targeted Re-Decoding for Quantized Small Language Model Reasoning


3. PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning


4. Reading and Steering Representations of Materials-Science Mechanisms in an Open-Weight Language Model


5. EvoDRC: A Self-Evolving Agentic Framework for Automated DRC Violation Repair


6. SenWorld: A Digital-Twin Simulation for Generating Context-Rich Evaluation Data


7. MOF-Sleuth: Tool-Grounded Reward Alignment for Explainable Fine-Grained MOF CIF Auditing


8. Rewarding Better Thinking for LLM Preference Alignment


9. Symbol and Footprint Database for Electronic Components by Agentic Recognition and Generation


10. Knowledge-Centric Self-Improvement


11. FORCE-Bench: A Benchmark, Dataset, and Evaluation Harness for Agentic AI in Enterprise Finance


12. HyGRL: Adaptive Hybrid Graph Reasoning for Multi-Entity Questions


13. CrackedPDFs: A Controlled Benchmark for Hidden Prompt Injection in PDFs


14. Mitigating Scaffolding Collapse in Socratic Tutors via Representation Alignment


15. Spectral-LSH: Sub-Quadratic Prompt Compression via Krylov-Projected Locality-Sensitive Hashing


16. Rethinking Uncertainty Evaluation in Large Language Models


17. Geometry-Guided Constraint Learning for LLM Safety Classification


18. Logic-Guided Data Extraction with Answer Set Programming and Large Language Models


19. Statistically Grounded Sparse-Feature Interventions for Activation-Space Control in Large Language Models


20. Lifted Representation Hypothesis in Language Models


21. Profile-Graph Memory for LLM Agents: Implicit Cross-Entity Traversal through Narrative Profiles


22. NEXUS: Structured Runtime Safety for Tool-Using LLM Agents


23. Information Discernment in Large Language Models


24. Benchmarking Confidential GPU Inference on NVIDIA H100 under Intel TDX


25. OpenEvoShield: Dual Non-Stationary Continual Defense for Open-World Multi-Agent System Attacks


26. FineServe: A Fine-Grained Dataset and Characterization of Global LLM Serving Workloads


27. Sound Probabilistic Safety Bounds for Large Language Models


28. The Maskability Index: Predicting Task-Objective Alignment in Pretrained Language Models


29. The Ethics of Autonomous AI Agents for Offensive Security


30. Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering


31. On the Systematic Challenges of Culturally Loaded Machine Translation: Dream of the Red Chamber as the Cultural Lens


32. Small, Free, and Effective: Orchestrating Open-Weight Small Language Models to Outperform Single LLM for Malware Analysis


33. Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning


34. SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD


35. ENTRAP-VL: A Taxonomic Probe for Dual Contextual Entrainment in Vision-Language Models


36. Reinforcement Learning for Large Language Model Selective Evidence Adoption from Contaminated Retrieval Results


37. Co-Evolving LLM Evaluators and Policies via DynamicRubric


38. Language-Specific versus Cross-Lingual Knowledge Graphs for Implicit Aspect Identification in Arabic: A Comparative Study of Reasoning and Adaptation Strategies


39. TINY_SCHILLER: A Drop-In German Drama Corpus for Small Language Models


40. When Shippers Become Algorithms: Candidate Exposure, Information Design, and the Concentration of LLM-Mediated Freight Markets


41. HijackKV: New Threat in Position-Independent KV Cache Reuse


42. Defense Against LLM Backdoors using Critical Neuron Isolation Pruning


43. Memory-Augmented Multimodal Large Language Models for Small Object Understanding in Streaming Aerial Videos


44. Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering


45. Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language Models


46. Beyond Fail-to-Pass: Iterative Hardening of Co-Generated Bug Reproduction Tests and Fixes


47. OPIUM: Mitigating Steering Externalities and Over-Refusal via Dual Objective Latent Optimization


48. An Isotropy-Preserving Spectral Cap for Muon: Theory and Three Case Studies


49. An Automated Framework for Extracting Reachable Attack Chains from Cyber Threat Intelligence Reports


50. Personalized Recommendation Tool Learning via Autonomous Language Agents


51. Reference-Free Evaluation of Reasoning in Open-Ended Question Answering


52. FedLSG: LLM-Enhanced Semantic Calibration for Federated Graph Backdoor Defense


53. PerfAgent: Profiler-Guided Iterative Refinement for Repository-Level Code Optimization


54. Adaptive Capitulation: A Structural Failure Mode of LLM Responses in Vulnerability Contexts


55. Causal dictionary learning reveals and validates transcription-factor binding features in genomic language models


56. Fine-grained Computation-Communication Overlap via Tile-level Signaling and Scheduling for Mixture-of-Experts


57. D3VL: Understanding Driving Scenes from 3D Time Series Data and Video with Language Models


58. Hybrid LLM-Guided Search for Quantum Reservoir Architecture Design


59. Integrity of peer-to-peer distributed LLM inference under malicious nodes


60. REGEN: Replay-recycling for Expert-to-Generalist distillation with Offline Reinforcement Learning


61. Guardrails as Scapegoats: Auditing Unfaithful Safety Refusals in Tool-Augmented LLM Agents


62. BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators


63. ChannelGuard: Safe Models Do Not Compose into Safe Multi-Agent Systems


64. JailMeter: An Evidence-Based Evaluation Framework for Jailbreak Attacks on Large Language Models


65. Auditing Retrieval-Augmented LLM Hypotheses for Longitudinal Cell Painting Morphology


66. From Trajectories to Prefixes: Reusing Teacher Trajectories via Replayed Prefixes and Online Continuation


67. Economic Evaluations of Language Models


68. Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework