Inside-Out: Measuring Generalization in Vision Transformers Through Inner Workings
Method for assessing model generalization in Vision Transformers via internal representations under distribution shift.
Method for assessing model generalization in Vision Transformers via internal representations under distribution shift.
Examines vulnerabilities in machine unlearning methods by analyzing internal representations and concept reintroduction.
DMax enables efficient parallel decoding in diffusion language models through progressive self-refinement.
Architecture combining frozen LLMs as nodes communicating through learned projections in a shared latent space.
Continual learning method using complementary self-supervised embeddings to improve replay buffer sample selection.
Parameter-efficient fine-tuning compression framework reducing communication costs for model adaptation.
Self-supervised pre-training method for time series classification with adaptive input handling.
Data augmentation method using adversarial training for out-of-distribution generalization on graphs.
KV cache offloading technique to reduce memory and latency overhead for long-context LLM inference.
Hybrid post-training combining reinforcement learning and distillation to improve LLM confidence calibration.
Test-time variational synthesis method for reinforcement learning in domains without verifiable rewards.
Impact of quantization on federated learning accuracy-efficiency trade-offs for aerospace predictive maintenance.
Analysis of how embedding dimensionality affects stability of graph node embeddings.
Mechanistic study of how steering vectors modify LLM behavior for alignment and refusal control.
Multi-agent system for language-agnostic code translation and validation across programming languages.
Framework for adaptive edge AI systems that adjust models during deployment as conditions change.
Memory architecture for efficient LLM inference on edge NPUs with optimized DRAM refresh for KV caches.
Benchmark dataset and evaluation for multimodal LLMs in manufacturing scenarios.
Industrial generative reranking system combining causality and utility for video search at scale.
Open-source framework for evaluating physical reservoir computing systems across various substrates.
LLM-based coding agents formalized 85K lines of topology proofs in Isabelle/HOL using ChatGPT and Claude.
Paper on generative reward models for LLM alignment using consistency-aware self-training to improve scalability.
Semi-autonomous multi-agent system for small molecule drug discovery using multi-modal AI agents and GNNs trained on 800M molecules.
RL-driven compiler using Soft Actor-Critic to jointly optimize ASIC architecture, memory hierarchy, and workload partitioning for on-device AI inference across technology nodes.
Framework using LLMs as semantic judges to validate and restructure outputs from unsupervised text clustering methods, improving coherence and grounding without labeled data.
CAMO is an ensemble technique for imbalanced text classification that optimizes minority class performance through hierarchical voting, confidence calibration, and uncertainty estimation.
Framework for understanding systematic variation in human-labeled training data, distinguishing between ambiguous items, divergent interpretations, and mistakes rather than treating all disagreement as noise.
Blink is an LLM serving architecture that removes the host CPU from the critical path by delegating orchestration and token control to GPU and SmartNIC, improving inference performance and datacenter resource utilization.
DIVERSED: relaxed speculative decoding for LLM inference using dynamic ensemble verification to improve token acceptance rates.
IatroBench: pre-registered study documenting how AI safety measures can cause harmful model behavior changes in medical advice contexts.
Dataset selection strategies for continual adaptation of generative recommenders under temporal distributional drift.
Methods to mitigate distribution sharpening in math RLVR through hint synthesis and annealing strategies.
Symbiotic-MoE: unified pre-training framework enabling Large Multimodal Models to generate images while maintaining understanding capabilities.
LSLoRA: investigation of sensitivity-positional co-localization in GQA transformers, restricting LoRA to optimally sensitive layers.
Graph learning framework for 3D engineering AI applications including CAE and CFD predictions with explainability.
SEARL: framework for self-evolving AI agents that jointly optimize policy and tool graphs to learn from trajectories without large-scale LLMs.
GRASS: gradient-based method for memory-efficient LLM fine-tuning using adaptive layer-wise importance sampling, balancing efficiency with model expressiveness.
Pipeline converting healthcare policy documents to executable BPMN models using LLMs for policy simulation and evaluation.
Recurrent-depth transformers enabling iterative reasoning to improve multi-hop knowledge composition in language models.
Learns first-order rules from image data without labels, automatically inventing predicates for explainable AI and enhancing LLM reasoning.
DACS mechanism for multi-agent LLM orchestration isolating per-agent context via registry and focus session modes to prevent context pollution.
GSSA-ViT framework using 3D Gaussian splatting for arbitrary-resolution weather forecasting and downscaling of atmospheric fields.
Unified framework viewing LLM post-training methods (SFT, preference optimization, RL, distillation) through off-policy and on-policy learning perspectives.
Studies data mixing strategies for LLM training, questioning domain definitions, human-model alignment, and impact of domain weighting on generalization.
Discusses risks of LLM-generated peer reviews and automated editorial processes, proposing RAG-XAI detection framework for identifying machine-generated content.
DSCA method for lifelong vision-language model editing via dynamic subspace concept alignment, preventing degradation from sequential edits.
Decomposes long-context reasoning in LLMs into atomic skills, automatically identifying and improving fundamental capabilities for complex reasoning.
First comprehensive survey of abductive reasoning in LLMs, defining taxonomy and exploring inference of plausible explanations from observations.
PrivFedTalk privacy-aware federated framework for personalized talking-head generation using diffusion models with identity-stable adapters.
LINE uses LLMs iteratively to explain individual neuron concepts in vision models without predefined vocabularies, enabling interpretability of neural networks.