GAGA: Gaussianity-Aware Gaussian Approximation for Efficient 3D Molecular Generation
Efficient 3D molecular generation method (GAGA) improving Gaussian probability path generative models by reducing sampling steps.
Efficient 3D molecular generation method (GAGA) improving Gaussian probability path generative models by reducing sampling steps.
Large-scale multilingual study (N=15,000) showing humans have more preference variation than LLMs; introduces Community Alignment Dataset for cultural pluralism.
Unsupervised learning approach to decompose neural network representation spaces into interpretable subspaces for mechanistic interpretability.
Automatic selection method for multimodal out-of-distribution detectors addressing robustness across diverse distribution shifts in video, audio, and sensor data.
Analysis of classification errors in automated audio/video processing for child development research and their downstream effects on findings.
Theoretical proof that GRPO RL algorithm with outcome reward models is equivalent to process reward models with Monte-Carlo-based weighting.
Framework for continual model merging addressing scalability of task vectors and functional drift through pre-, during-, and post-merging interventions.
Comparative study of scaling laws for xLSTM versus Transformers, showing competitive performance with linear time-complexity for LLMs at billion-parameter scale.
Low-rank adapter approach (SuLoRA) for handling subject-specific distribution shifts in EEG brain decoding by decomposing weights into shared and subject-specific components.
Theoretical work identifying non-standard vector spaces where neural networks act as linear operators using transport of structure from algebra.
Method for transferring task vectors across different pre-trained foundation model versions using gradient-sign masking to handle parameter space misalignment.
TraDy transfer learning scheme for memory-efficient fine-tuning using architecture-dependent layer importance and dynamic channel selection.
FATE benchmark for formal algebra theorem proving with multiple difficulty levels to evaluate LLM capabilities beyond contest problems.
ExPairT-LLM selects best generated code via pairwise queries without assuming LLM correctness on all inputs.
Probabilistic certification framework for SmoothLLM defense against jailbreaking with relaxed assumptions for practical LLM safety.
MIST proposes neural network-based mutual information estimator trained on synthetic distributions for variable-size inputs.
CDLM accelerates diffusion language models through consistency modeling and enables KV caching for faster parallel generation.
Conductor model uses reinforcement learning to discover coordination strategies and communication topologies among specialized LLMs.
Transfer learning approach for discrete diffusion models in small-data regimes using classifier ratio-based guidance adapted from continuous models.
Policy optimization approach for molecular design that learns amortized policies transferable across unseen molecular structures.
Framework addressing behavioral staleness in asynchronous federated learning to improve training performance when clients operate at different speeds.
Analysis of how neural network capabilities emerge during training across model scales, tracking representation collapse and reorganization across 120+ emergence events.
New optimizer combining Adam's adaptive moments with Muon's orthogonalized momentum for improved LLM training efficiency and performance.
Method for optimally allocating observations between explainable and black-box models to maximize ensemble performance while maintaining interpretability guarantees.
Research on organizational governance frameworks for generative AI and LLMs, addressing technical and business perspectives on risk and opportunity management.
Proposes fine-tuning LLMs with bags of sentences for improved topic modeling over classical LDA and out-of-the-box pretrained encoders.
Proposes contextual guidance approach for knowledge distillation enabling smaller LLMs to provide coherent multi-turn responses in customer interactions.
CausalBGM applies Bayesian generative modeling with AI for causal inference in observational studies with high-dimensional covariates.
CAIMAN framework uses causal action influence detection in reinforcement learning for sample-efficient legged robot loco-manipulation tasks.
Integrates physical quantities into deep generative models for solar magnetic active region generation and retrieval with scientific interpretability.
Studies impact of missing data mechanisms on algorithmic fairness, showing demographic-linked missingness introduces bias in ML systems.
CAE repurposes critic networks in deep RL as exploration drivers using multi-armed bandit techniques without additional parameters.
ConformalNL2LTL translates natural language instructions to linear temporal logic formulas with conformal correctness guarantees for autonomous systems.
Addresses learning spreading dynamics in social networks with hidden individual statuses using classification methods with observable intermediate indicators.
Presents AstroSage-Llama-3.1-70B, a domain-specialized 70B LLM for astronomy Q&A and research assistance exceeding general-purpose models.
Introduces V²-VLNCE benchmark and view-invariant post-training framework for vision-language navigation in embodied AI agents.
Studies eigenvalue behavior of perturbed random matrices relevant to DNN weight matrices and pruning techniques based on random matrix theory.
Uses vision-language models and graph neural networks to detect deepfakes with textual explanations for improved robustness and generalization.
Proposes matrix-based dictionary learning for transformer weight sharing to reduce computational and memory demands of LLMs by exploiting inter-block redundancy.
Two-player Markov game framework studying human-AI interaction for balancing agent autonomy and safety through minimal control interface.
Kubernetes scheduler using LLM to interpret natural language hints for semantic, intent-driven cluster workload allocation with soft affinity preferences.
Controlled study of how pretraining discourse about AI behavior influences LLM alignment outcomes, with 6.9B-parameter model experiments on behavioral priors.
Analysis of LLM benchmark saturation and the challenge of creating discriminative tasks as frontier models improve, discussing feasibility of future benchmarking.
RAG-based framework for question answering over multi-hour audio with temporal grounding, addressing context-length limitations of audio-language models.
CoreCraft RL environment from EnterpriseBench suite for training generalizable agents on high-fidelity enterprise customer support simulations with 2,500+ entities and 23 tools.
Multi-agent LLM and vision framework for robotic manipulation with closed-loop feedback, enabling task planning without fine-tuning in dynamic environments.
Benchmark for evaluating LLM planning and reasoning by navigating Wikipedia hyperlinks to reach target pages, testing world knowledge and look-ahead planning across multiple model variants.
Aqua is a CLI message tool designed for AI agents. Title only, minimal technical details provided.
React portfolio that dynamically re-architects its DOM based on LLM intent analysis using Llama-3 via Groq, adapting content for different audiences (recruiters, founders, engineers).
Case study documenting 5 failure modes from running AI agents autonomously: auto-rotation loss, documentation trap, market inefficiency, static models, and monitoring gaps.