Benchmark for evaluating multimodal LLMs on handwritten STEM student solutions with mathematical formulas and diagrams, addressing authentic domain-specific evaluation gaps.
TernaryLM: Language model trained natively with 1.5-bit quantization achieving memory-efficient deployment on edge devices while maintaining language modeling capability.
Benchmark evaluating LLM-based coding agents on their ability to learn from context and reuse experience across related software engineering tasks in repositories.
Comparative study of CNN architectures (VGG, ResNet, GoogLeNet) analyzing relationship between depth and trainability in image recognition.
DUET-VLM: dual-stage token reduction framework for vision-language models reducing computational cost while maintaining accuracy during training and inference.
PedaCo-Gen: pedagogically-informed human-AI system for collaborative instructional video generation using Cognitive Theory of Multimedia Learning.
Layer gradient analysis method for identifying optimal layers in LLMs for knowledge editing while preserving model behavior on unrelated inputs.
SpotIt+: open-source verification tool for Text-to-SQL evaluation using bounded equivalence checking and constraint-mining for practical query discrepancies.
DiFlowDubber: two-stage approach for automated video dubbing using discrete flow matching for expressive prosody and precise audio-visual synchronization.
AgentTrace: lightweight framework for post-hoc root cause analysis in deployed multi-agent systems using causal graph tracing from execution logs.
Study showing LLMs struggle with private library code generation despite API documentation; proposes teaching methods for private-library-oriented code generation.
Analysis of multimodal LLMs generating natural language explanations for face verification decisions on unconstrained images.
Goedel-Code-Prover: hierarchical proof search framework for automated code verification in Lean 4 using LLMs to decompose complex verification goals.
Analysis of how AI scaling laws reshape classical Amdahl's Law for modern heterogeneous computer architectures with specialized accelerators and tensor datapaths.
KG-Hopper: reinforcement learning framework enabling compact open-source LLMs to perform knowledge graph reasoning for multi-hop KBQA tasks.
mSFT: iterative algorithm for multi-task supervised fine-tuning that addresses heterogeneous overfitting by dynamically adjusting compute budget across datasets.
KALAVAI: quantitative model predicting when independently trained specialist LLMs can be fused post-hoc with measurable performance gains; includes practical prediction formula.
EVA: reinforcement learning framework for video agents using MLLMs with adaptive reasoning to handle long video sequences and temporal dependencies efficiently.
MDKeyChunker: structure-aware chunking pipeline for Markdown documents with single-call LLM enrichment to improve RAG accuracy and reduce metadata extraction overhead.
arXiv paper analyzing response homogenization in RLHF-aligned LLMs and its effects on uncertainty estimation methods.
arXiv paper introducing scalability coefficients for detecting problematic items in large-scale AI benchmarks using isotonic regression.
arXiv paper demonstrating dual-layer side-channel attacks on local Vision-Language Models exploiting dynamic preprocessing vulnerabilities.
MAGNET: decentralized system for autonomous generation and training of domain-expert language models using autoresearch and BitNet ternary quantization.
Theoretical analysis of simplicity bias in neural networks using minimum description length principle and compression framework.
Investigation of whether LLMs perform genuine in-context molecular property prediction or rely on memorization despite potential training data contamination.
Analysis of activation-based probes for detecting misaligned AI systems, showing blind spots in detecting coherent misalignment versus deception.
DRiffusion: parallel sampling framework accelerating diffusion model inference through draft-and-refine process with skip transitions.
EngineAD real-world multivariate anomaly detection dataset from vehicle fleet sensor telemetry with expert annotations for safety-critical domain.
ARTA joint training framework for adversarially robust multivariate time-series anomaly detection using min-max optimization and information retention.
Somax composable Optax-native stack for second-order curvature-aware training with modular APIs for operators, estimators, and preconditioners.
QuitoBench open benchmark for time series forecasting covering eight trend-seasonality-forecastability regimes with regime-balanced dataset design.
GLU framework for sparse spatiotemporal reconstruction and forecasting using global-local-uncertainty fusion with unified state representation.
H-Node ANC mechanistic framework identifies and defends hallucination representations in transformer LLMs at individual hidden-state dimensions.
Study of LLMs' theory of mind capabilities using behavior-based testing to assess their ability to self-model and model other agents.
Dynamic Tokenization via Reinforcement Patching learns variable-sized data-driven patches for long-horizon sequence models with zero-shot transfer capability.
Framework assessing robustness of LLM-enhanced Graph Neural Networks against poisoning attacks targeting both graph structure and textual attributes.
TinyML pipeline for real-time acoustic anomaly detection on IoT microcontrollers for environmental sound monitoring without cloud processing.
PEANUT proposes perturbation-based adversarial attack method targeting robustness vulnerabilities in Graph Neural Networks through topology modifications.
PruneFuse strategy uses pruned networks for efficient data selection and fuses them with original networks to optimize deep neural network training.
Complexity analysis of optimal graph rewiring to address oversmoothing and oversquashing in deep graph neural networks.
Unified framework for data-centric dynamic training of LLMs with consistent interfaces for data selection and reweighting.
Study of LLM-based AI scientist agents learning from iterative experimental feedback in cell screening with 800 replicated experiments.
Analysis of optimization trade-offs in asynchronous federated learning addressing gradient staleness and client bias.
Knowledge distillation approach for deploying Transformer-based reinforcement learning on resource-constrained energy management devices.
Formal framework for measuring uncertainty in LLM text generation accounting for prompting, generation, and interpretation stages.
Survey of generative modeling in protein design covering neural representations, conditional generation, and evaluation standards.
Neuro-symbolic approach combining neural networks and domain knowledge for process anomaly detection from event logs.
Koopman autoencoder-based least-squares policy iteration algorithm enabling automatic feature learning in reinforcement learning.
Framework using Shapley values to measure and explain unfairness in machine learning models under group fairness criteria with inference methods.
SPECTRA: spectral-informed neural network for sensor-based activity recognition optimized for edge deployment with low latency and privacy.