A Contextual Help Browser Extension to Assist Digital Illiterate Internet Users
Browser extension combining curated dictionary with OpenAI LLM to provide contextual help for technical terms via tooltips.
Browser extension combining curated dictionary with OpenAI LLM to provide contextual help for technical terms via tooltips.
Neural network training robust to both label noise and adversarial attacks via unified loss function approach.
Benchmarking framework for reinforcement learning algorithms using stochastic converse optimality with known optimal policies.
Evolutionary algorithms to automatically design efficient multigrid cycles for solving partial differential equations.
Cross-domain few-shot learning with vision-language models like CLIP for fine-grained visual recognition tasks such as medical diagnosis.
Benchmark and analysis of hallucinations in multimodal LLMs with fine-grained negative queries covering multi-object, multi-attribute, and multi-relation scenarios.
Post-training framework for small local LLM agents performing Linux privilege escalation tasks with verifiable rewards and resource constraints.
Adaptive guidance method for retrieval-augmented masked diffusion models that resolves conflicts between retrieved context and parametric knowledge.
Benchmark for evaluating vision-language model reasoning and segmentation performance under adverse weather conditions with degraded visual cues.
Framework for testing LLM trading agents with anonymized market data to validate genuine market understanding versus memorized ticker associations.
Evaluation of Segment Anything Model 3 for eye image segmentation comparing text and visual prompting modes against prior SAM versions.
Survey of machine learning methods for network intrusion detection and adversarial learning techniques for synthetic attack data generation.
Training-free fine-grained visual recognition method using sample-wise adaptive reasoning with large vision-language models for subordinate-level category classification.
Analysis showing attention sinks in transformers induce gradient concentration during backpropagation under causal masking, affecting training dynamics.
LLM training approach using generator-verifier co-evolution to escape consensus trap and improve reasoning without ground-truth labels via reinforcement learning.
Theoretical analysis of sparsity in infinite-width shallow ReLU networks trained with total variation regularization using duality theory.
On-model anomaly detection method that leverages primary model representations to detect distributional shifts without separate AD models.
Video world model framework with inverse dynamics rewards that ensures generated robot action sequences satisfy rigid-body and kinematic constraints.
Post-training quantization method for large vision-language models using quantization-aware integrated gradients to reduce memory and computational overhead.
Study of dropout-induced variability and uncertainty in transformer models across 19 architectures using Monte Carlo sampling for inference-time evaluation.
Multimodal framework for automated program repair using LLMs that jointly reasons over code, issue descriptions, and GUI screenshots to improve debugging workflows.
Reinforcement learning method for training code search agents to localize relevant files, classes, and functions in large repositories as prerequisite for coding tasks.
arXiv paper on inferring stage-play spatial layouts from narrative text, demonstrating language model spatial reasoning for automating dramaturgy.
arXiv paper proposing GeCO, time-unconditional flow matching framework for adaptive robotic control using diffusion models.
arXiv paper investigating how LLMs compute verbal confidence scores and whether they're generated just-in-time or cached during inference.
arXiv paper introducing DiscoGen, procedural generator for algorithm discovery tasks addressing evaluation and contamination issues in ML benchmarks.
arXiv paper proposing domain-grounded tiered retrieval architecture to reduce LLM hallucinations through systematic factual verification.
Reinforcement learning framework for adaptive mixed-precision quantization of LLMs on resource-constrained devices with per-layer bit width optimization.
LLM-based linter detecting methodology bugs in scientific Python code that produce plausible but incorrect results.
Analysis of privacy risks from enterprise data exposure in LLM-integrated systems with optimal differential privacy tradeoffs for agents.
First systematic safety evaluation of LLMs across 12 Indic languages using 6,000 culturally grounded prompts in low-resource settings.
Technique for converting grouped-query attention to multi-head latent attention with improved expressivity and reduced KV-cache costs.
Method for efficient long-form video processing with hierarchical grid representation enabling lossless, scalable video understanding.
Open-source tool and benchmark for reducing code regressions in AI coding agents using AST-based impact analysis and test-driven development.
Automated approach for generating repository-level vulnerability detection datasets at scale beyond function-centric benchmarks.
Framework augmenting 2D vision-language models with 3D spatial understanding and viewpoint-aware reasoning capabilities.
Token pruning approach for efficient video vision-language models addressing temporal redundancy in video-based downstream tasks.
Study of polysemanticity in language models using sparse autoencoders to identify semantic interference patterns transferable across models.
Framework for explainable depression diagnosis using multimodal large language models on interview videos with depression score predictions.
Evaluation showing large multimodal models struggle with inductive physical reasoning tasks beyond laws observed during training.
Benchmark and analysis of multimodal agents' GUI interaction capabilities, identifying toggle control as a key bottleneck in graphical user interface automation.
Multi-agent calendar assistant using graph-structured coordination with supervisory agent overseeing specialized task agents for natural language Google Calendar management.
Diagnostic evaluation of LLM reasoning capabilities through implicit causal chain discovery task, testing nine LLMs on mechanistic causal reasoning in climate discourse.
Multi-agent LLM system for longitudinal psychological counseling with emotional understanding, adaptive strategies, and long-term memory capabilities.
Research on improving LLM safety evaluation using multi-agent debate with HAJailBench, a 11,100-sample human-annotated jailbreak benchmark across diverse attack methods.
Safety-preserving post-training quantization via contrastive alignment loss to maintain behavioral safety during LLM model compression.
Study on steering LLM probabilistic beliefs under informative missingness patterns for improved clinical reasoning with incomplete data.
Stepwise Think-Critique framework unifying reasoning and verification in LLMs through intertwined critical thinking for robust problem-solving.
CircuitLM multi-agent LLM pipeline generating circuit schematics from natural language, addressing hallucination and constraint violation issues in EDA.
PaperScout autonomous agent for academic paper search using reinforcement learning to dynamically decide tool invocation for complex conditional queries.