Fairness Evaluation and Inference Level Mitigation in LLMs
Inference-level mitigation techniques for fairness issues in LLMs without expensive retraining.
Inference-level mitigation techniques for fairness issues in LLMs without expensive retraining.
RAG framework using hierarchical sequences for multi-hop reasoning over heterogeneous sources (documents, tables, KGs).
Uncertainty estimation method for surgical VQA systems accounting for question context in safety-critical applications.
Continual learning framework for binary malware detection with controlled forgetting to adapt as threats evolve.
Accelerates visual autoregressive generation via partial verification skipping in speculative decoding.
Studies why LLM agents disclose information to external parties contrary to user instructions, examining alignment behavior.
Benchmark for evaluating LLM and vision-language model comprehension of musical scores in ABC and PDF formats.
Study identifying and characterizing biases in machine-generated text detection systems across different text types and domains.
Framework abstracting LLM reasoning traces into functional steps using Schoenfeld's theory to analyze mathematical reasoning structure.
Multimodal fusion approach using Fisher information for vulnerability detection in code combining sequences and property graphs.
Analysis showing reasoning distillation via supervised fine-tuning fails to transmit cognitive structure of reinforcement-trained reasoning models.
Framework reweighting negative samples during LLM fine-tuning to improve data efficiency beyond supervised fine-tuning and rejection sampling.
Parameter-efficient adaptation method preserving geometric structure of pre-trained models during reinforcement learning with verifiable rewards.
NPU architecture design optimized for diffusion-based LLM inference with bidirectional attention and block-wise KV cache patterns.
Projection model predicting sensorimotor norms from word embeddings to ground language understanding in embodied experience.
Framework replacing binary preference labels with continuous utility scores for fine-grained alignment of LLM reasoning capabilities.
Framework leveraging LLMs for solving parametric PDEs through code generation and operator inference techniques.
Multi-agent reinforcement learning approach for overlay multicast routing with network situational awareness.
Evaluation of adversarial safety datasets revealing they rely on triggering cues rather than reflecting real-world attacks.
Methods for dynamic rollout allocation and advantage modulation in reinforcement learning with verifiable rewards for LLM reasoning.
LLM-driven framework combining threat modeling and formal verification for SoC security property generation and assertion checking.
Framework distinguishing cognitive amplification from cognitive delegation in human-AI systems with metrics for hybrid performance.
Study of contextual bias and framing effects in LLM-based code review systems for vulnerability detection in CI/CD pipelines.
Benchmark studying comic-template jailbreaks that exploit visual narratives to compromise multimodal LLM safety alignment.
System using LLMs to jointly rank cited papers within a citing paper to measure relative impact with positional bias mitigation.
Unlearning framework for vision-language-action embodied foundation models to remove unsafe or privacy-sensitive behaviors from robotic policies.
Fast visual navigation approach using rectified Schrödinger bridge matching for embodied AI agents with few integration steps.
Method for improving LLM safety alignment across low-resource languages by addressing semantic bottleneck and language-agnostic understanding.
Framework for evaluating automated repair of logical vulnerabilities in software using recent LLM techniques for semantic understanding.
Optimization technique leveraging mixture-of-experts elasticity for self-speculative decoding in LLM serving to reduce memory bottlenecks.
Study of chain-of-thought prompting with LLMs for code deobfuscation, using step-by-step reasoning for control flow analysis.
Analysis of agentic AI systems and LLMs in software engineering, discussing AI threats to SE tasks like testing, bug fixing, and integration work.
Technique for reducing time-to-first-token in LLM inference by overlapping context streaming and prefill under concurrent requests.
Method for verifying reasoning traces in diffusion language models using geometric perspective and bidirectional consistency checking.
Benchmark for evaluating biases in multimodal LLMs used as automatic judges, identifying failures in integrating visual and textual evidence.
Benchmark for evaluating LLM agents on threat hunting in security operations, using 106 real attack procedures from Windows event logs across MITRE ATT&CK techniques.
Comparative analysis of output consistency across GPT-4, Claude, and Gemini for repeated exercise prescription generation in clinical scenarios.
Benchmarking study measuring biological misuse capability across advanced LLMs using STEM prompts to assess AI safety safeguards.
Studies visual communication modalities for mobile GUI agents, balancing user transparency and multitasking through adaptive display approaches.
Surrogate modeling framework using simplified models to interpret and explain black-box LLM behavior in medical prediction tasks.
Interactive framework enabling industrial robot skill adaptation through kinesthetic, natural language, and graphical interaction modalities.
Proposes structured memory units for LLMs that store knowledge separately from parameters, enabling efficient updates without retraining.
Systematic study of prompt optimization for LLM-as-a-Judge evaluations in legal question answering, examining transfer across different judge models.
Research integrating working memory constraints into Transformers via attention variants, trained on limited data for improved language performance.
Surrogate modeling framework using simplified models to interpret and explain black-box LLM behavior in medical prediction tasks.
Interactive framework enabling industrial robot skill adaptation through kinesthetic, natural language, and graphical interaction modalities.
Research integrating working memory constraints into Transformers via attention variants, trained on limited data for improved language performance.
SF startup Andon Labs deployed AI agent to manage physical retail store; reports of operational issues including overordering and wage disparities.
DeepSeek V4 release with Flash (284B/13B active) and Pro (1.6T/49B active) models, 1M token context, competitive on coding and reasoning tasks.
Cost analysis estimates Nvidia B200 GPU production at $5,700-$7,300, with HBM memory and packaging comprising two-thirds of costs.