Large Language Models for Wireless Communications: From Adaptation to Autonomy
Survey of LLM applications in wireless communications covering adaptation, autonomy, and intelligent system design for complex communication networks.
Survey of LLM applications in wireless communications covering adaptation, autonomy, and intelligent system design for complex communication networks.
Multi-agent pipeline for street design generation combining image generation and infrastructure design for urban planning visualization.
Hilbert system combines informal LLM reasoning with formal theorem proving in Lean 4 for verifiable mathematical proofs.
ReasoningBank memory framework enables LLM agents to learn from interaction history and distill generalizable reasoning strategies for continuous tasks.
Zephyrus agentic framework combines weather foundation models with LLM reasoning for interactive scientific workflows in meteorology.
Study showing AI agents fail under realistic user behavior variations; proposes high-fidelity human trait simulations for robust agent testing.
PREFINE framework enables personalized story generation using simulated user critics and rubric generation without explicit user feedback.
Multi-agent debate framework using small language models for cost-efficient LLM safety evaluation, with HAJailBench benchmark for jailbreak testing.
Alignment-Aware Quantization: PTQ method for efficient LLM deployment that preserves behavioral alignment and safety properties, not just minimizing reconstruction error.
SpatialBench: benchmark measuring multimodal LLM spatial cognition across hierarchical abilities for real-world physical environment interaction.
Analysis of multi-agent path finding algorithm design trade-offs under realistic robot execution constraints for warehouse and manufacturing applications.
Stepwise Think-Critique: unified framework integrating reasoning and verification in LLMs for robust, interpretable problem-solving with intertwined evaluation.
FusionRoute: token-level collaboration method enabling multiple specialized LLMs to work together, combining domain expertise efficiency with generalization.
VisTIRA: tool integration approach addressing modality gap where VLMs underperform on visual math problems compared to text-based versions.
LogicSkills: benchmark isolating three fundamental logical reasoning skills in LLMs: formal symbolization, countermodel construction, and logical inference.
Empirical study of latent chain-of-thought in LLMs using structural causal models to analyze intermediate computation steps beyond correlation-based probes.
Benchmark comparing zero-shot Time Series Foundation Models against classical methods for annual institutional demand forecasting under data sparsity.
MemPO: self-memory policy optimization approach enabling long-horizon agents to proactively manage memory content aligned with task objectives.
Model Medicine: framework for understanding, diagnosing, and treating AI model disorders using biological organism principles as analogy for model analysis.
UIS-Digger: LLM-based research agent system for unindexed information seeking, addressing blind spots where vital information isn't captured by search engines.
Evaluation of frontier AI models' autonomous cyber-attack capabilities on multi-step scenarios, tracking capability trends across 18 months of model releases.
Framework for automated skill acquisition in modular AI agents through mining open-source repositories to extract procedural knowledge and specialized expertise.
Framework for unlearning relational safety failures in multimodal LLMs where combinations of benign concepts become unsafe when linked by specific relations.
Gradient Atoms: unsupervised method for discovering and attributing model behaviors via sparse decomposition of training gradients without requiring predefined queries.
OpenHospital: interactive benchmark arena for evolving and evaluating LLM-based collective intelligence systems using physician and patient agents.
Theoretical analysis arguing that LLM's most valuable capabilities are the unexplainable components that cannot be captured by discrete rule systems.
SAGE framework: multi-agent reinforcement learning system for improving LLM reasoning without large human-labeled datasets, using self-play and closed-loop feedback.
LLAMAFUZZ uses LLMs to enhance greybox fuzzing for structured data, improving mutation strategies beyond random approaches.
TS-Reasoner agent integrates LLM reasoning with domain-specific code for multi-step time series inference and automated analysis tasks.
Systematic literature review of LLM security benefits and drawbacks in code generation, vulnerability detection, and remediation tasks.
Interdisciplinary study on copyright law implications of training generative AI via web scraping, covering fair use and TDM exceptions.
MASS method merges multiple fine-tuned models via adaptive subspace selection, improving accuracy over existing merging approaches without retraining.
Research shows diverse AI personas in generative AI reduce homogenization in collaborative creative outputs compared to single-persona systems.
CRBench: Real-world benchmark for text-to-chart retrieval using synthesized semantic insights.
FALCON: Method addressing false negatives in vision-language pretraining through contrastive learning.
Systematic literature review of explanation user interfaces for interpretable AI systems.
BiomedSQL: Text-to-SQL benchmark for scientific reasoning over biomedical databases requiring domain knowledge.
VERINA benchmark for evaluating LLM code generation with joint code, specification, and proof generation.
Robust aggregation method for distributed learning systems defending against Byzantine attacks.
Differential privacy techniques for LLMs applied to radiology report classification tasks.
Benchmark evaluation of LLM effectiveness for text diacritization in Arabic and Yoruba with MultiDiac dataset.
Structured instruction approach to improve chart-to-code generation in multimodal LLMs with iterative refinement.
VideoITG: Frame selection method for efficient video understanding in video-LLMs using instructed temporal grounding.
Systematic study of LLM capabilities for discrete choice modeling with analysis of prompting strategies.
Method for detecting LLM confabulations using token-level uncertainty estimation for reliability in agentic applications.
Dynamic weighting approach integrating supervised fine-tuning and reinforcement learning for LLM post-training alignment.
Vision-based learning framework for omnidirectional bipedal locomotion using depth images on challenging terrain.
CodeGym: Reinforcement learning framework for training LLM agents to use tools generalizing across new tasks and workflows.
LANCE: Low-rank compression technique for reducing activation memory in on-device continual learning.
ERGO: Two-stage coarse-to-fine pipeline for efficient high-resolution image processing in vision-language models.