Generative Interfaces for Language Models
Proposes generative interfaces paradigm to move LLM interactions beyond linear request-response format for more efficient multi-turn, information-dense, and exploratory tasks.
Proposes generative interfaces paradigm to move LLM interactions beyond linear request-response format for more efficient multi-turn, information-dense, and exploratory tasks.
Study evaluates safety-aligned versus uncensored LLMs on hate speech detection, revealing trade-offs between model censoring and detection performance across political personas.
Method improves LLM factuality by constructing knowledge graphs at inference time instead of using unstructured text retrieval, enhancing reasoning and reducing irrelevant information influence.
ALIGNS: LLM-based approach for building nomological networks in psychological measurement to establish construct validity.
Method for generative recommenders learning decomposed contextual token representations combining pretrained and collaborative signals.
Theoretical analysis of watermarking robustness for generative models using zero-bit tamper-detection codes.
Study evaluating quantization impact on vision-language models across reliability metrics beyond accuracy for OOD detection.
CoSpaDi: training-free LLM compression using calibration-guided sparse dictionary learning as alternative to low-rank approximations.
Analysis of hybrid attention architectures combining local sliding window and global attention for improved long-term memory.
Ergodic risk measures framework for continual RL agents balancing knowledge retention and adaptation to new environments.
Framework for multi-source reasoning alignment in MLLMs addressing concept drift in non-stationary environments.
TokenChain: fully discrete speech chain coupling semantic-token ASR with two-stage TTS for joint improvement of speech tasks.
CleverCatch: knowledge-guided weak supervision model for healthcare fraud detection with limited labeled data.
TokenTiming: speculative decoding acceleration for LLM inference enabling draft and target models with different vocabularies.
Method training coding agents using compiler and language server feedback as supervision signals to improve program correctness.
Method using diffusion models for solving linear inverse problems by balancing noise integration to maintain generative quality.
Study on how understanding linguistic implicature improves LLM alignment with user intent in human-AI interaction.
TetraJet-v2: 4-bit quantized training method for LLMs using NVFP4 format with techniques to suppress oscillation and control outliers.
Method for protecting deep neural networks from unauthorized use by enabling model usage control without embedding access keys in parameters.
Synthetic benchmark dataset for video anomaly detection with improved scene diversity, balanced coverage, and temporal complexity evaluation.
Survey of intelligent agents with emotional intelligence capabilities for human-computer interaction and affective computing systems.
Balanced fine-tuning method for aligning LLMs with biomedical knowledge by addressing uncertainty structure differences in scientific text.
Physics-informed generative AI framework for thermal circuit analysis, accelerating early-stage design validation vs traditional FEM simulation.
2.4B parameter multilingual vision-language model achieving SOTA performance on multilingual VQA with token-efficient image processing.
Training-free algorithm for unbounded terrain generation using diffusion models, combining procedural noise efficiency with learned fidelity.
Neuro-symbolic explainability framework for public-sector AI systems, linking AI decisions to legal requirements using structured knowledge representation.
Cognitive-semantic analysis of prompt engineering as natural-language control mechanism, using frame activation and salience to explain LLM behavior modification.
Framework for multimodal fake news detection using conflict-consensus approach that leverages cross-modal discrepancies instead of enforcing consistency.
InstructMoLE: Parameter-efficient fine-tuning of Diffusion Transformers using Mixture of Low-rank Experts with instruction-guided routing for multi-conditional image generation.
Research using Large Vision-Language Models to align task-specific vision models with human domain knowledge, reducing spurious correlations.
Using large vision-language models to improve alignment of task-specific vision models with human knowledge.
OpenAI GPT-5 system card describing unified architecture with fast base model, deeper reasoning model, and real-time router for task-appropriate model selection.
RepoReason benchmark for evaluating agentic code reasoning at repository level with white-box diagnostics for logical consistency across interdependent file systems.
Study of emoticon semantic confusion vulnerability in LLMs where ASCII-based emoticons cause misinterpretation and unintended actions.
STAGE benchmark for evaluating model reasoning over full-screenplay narratives, testing story comprehension, character tracking, and multi-form generation consistency.
RoundTripCodeEval benchmark evaluating code-LLM reasoning consistency through forward-backward execution via lossless compression algorithm round-trip fidelity.
Empirical study testing whether professional translators can identify AI-generated text from ChatGPT-4o versus human authors without specialized training.
Method using vector alignment to resolve jailbreak-overrefusal trade-off in safety-aligned LLMs by separately encoding answer vectors and safety judgment vectors.
Diagnostic study comparing iterative RAG with static RAG for multi-hop scientific question answering, analyzing when synchronized retrieval-reasoning loops provide benefits over single-pass approaches.
TCLA framework enables cross-session neural decoding with limited target data using task-conditioned latent alignment from source sessions.
VERGE neurosymbolic framework combines LLMs with SMT solvers for verification-guided reasoning, decomposing outputs into formal logic for consistency checking.
CoFrGeNets introduces continued fraction-inspired architecture replacing attention and feed-forward layers in Transformers for improved language generation.
GRACE framework unifies quantization-aware training and knowledge distillation for efficient vision-language models using information bottleneck principle.
VideoGPA adds geometric priors to video diffusion models via self-supervised preference alignment to improve 3D structural consistency in generated videos.
MCP-Atlas large-scale benchmark for evaluating LLM tool-use competency with 36 real Model Context Protocol servers, capturing real-world tool invocation complexity.
STEP warm-starts diffusion-based visuomotor policies with spatiotemporal consistency prediction to reduce inference latency for real-time robotic control.
Control Reinforcement Learning trains policy to select sparse autoencoder features for interpretable token-level steering of LLMs, producing explainable intervention logs.
Comprehensive investigation of post-training pipeline for LLM-based vulnerability detection, demonstrating on-policy RL with GRPO outperforms supervised fine-tuning approaches.
CrispEdit scalable second-order algorithm for LLM editing that preserves general capabilities while making targeted behavior changes, treating capability preservation as explicit constraint.
HistCAD benchmark and dataset for parametric CAD generation with constraint preservation under edits, measuring design intent retention rather than just reconstruction fidelity.