Compiling Deterministic Structure into SLM Harnesses
Semantic Gradient Descent framework compiles agentic workflows into discrete execution plans (DAG topologies, prompts, code) for efficient small language model deployment.
Semantic Gradient Descent framework compiles agentic workflows into discrete execution plans (DAG topologies, prompts, code) for efficient small language model deployment.
arXiv research: Language models can detect, localize and verbalize activation perturbations like dropout and Gaussian noise.
Design challenges for safety-critical autonomous systems integrating AI components, covering safety, security, reliability, and certification across abstraction layers.
Schema-aware memory architecture for AI agents enabling exact fact storage, state tracking, updates, and structured retrieval beyond text embeddings.
D3-Gym dataset with 565 verifiable tasks from real scientific repositories for evaluating LLM and agent capabilities in data-driven discovery.
Research infrastructure mapping methodological evolution in AI as a graph structure to support AI-driven research agents accessing scientific knowledge.
Collaborative diffusion models for synthetic image generation addressing data availability, computational requirements, and privacy challenges.
Decision framework for systematic evaluation of bias and fairness in LLMs across different deployment contexts and use cases.
Algorithm for learning constrained MDPs using parameterized policies with entropy and quadratic regularizers, achieving last-iterate convergence.
Theoretical paper examining how large language models form internal representations, bridging disagreement between LLM optimists and pessimists.
Analysis of system 1 thinking capability in large reasoning models, evaluating intuitive responses with minimal tokens for efficiency and difficulty awareness.
Research on causal inference for autonomous mobile robots in dynamic environments, modeling human behavior and interactions.
Study comparing exploration-exploitation tradeoffs in LLMs versus humans using multi-armed bandit experiments for sequential decision-making.
Legal Assist AI: Lightweight domain-adapted framework for providing legal assistance in Indian context using efficient LLM fine-tuning.
ML-Agent: Framework reinforcing LLM agents for autonomous machine learning engineering, moving beyond prompt-based paradigms for better generalization.
Disentangled Safety Adapters: Framework decoupling safety computations from base models using lightweight adapters for efficient guardrails and flexible alignment.
VGR: Framework for visual grounded chain-of-thought reasoning that performs reasoning in visual space rather than pure language, for complex visual tasks.
Method for refining LLM-generated code through property-oriented, minimal feedback instead of test quantity, improving functional correctness.
Benchmark comparing multimodal foundation models (GPT-4o, Gemini, Claude, Qwen, Llama) on standard computer vision tasks beyond question answering.
ExCyTIn-Bench: First benchmark for evaluating LLM agents on autonomous cyber threat investigation using multi-hop evidence chains from security logs.
Study investigating intentional deception in LLMs on benign prompts without explicit hidden objectives, assessing trustworthiness for decision-making tasks.
InterChart benchmark evaluating vision-language models on reasoning across multiple related charts for scientific, financial, and policy analysis tasks.
LinkAnchor: Autonomous LLM agent for recovering issue-to-commit links in software repositories, addressing low linking rates on GitHub.
Framework for using LLMs to perform reasoning-intensive regression tasks like rubric-based scoring and dense reward modeling, beyond standard text analysis.
Qualitative study of how generative AI reshapes product development workflows through natural language-to-code translation, based on 22 interviews across teams.
Method using functional representations to trace evolutionary relationships between millions of LLMs, addressing undocumented fine-tuning and adaptation chains.
Benchmark framework (Game-Time) for evaluating temporal dynamics capabilities in conversational spoken language models, including timing and simultaneous speech handling.
Research on training vision-language models for geospatial reasoning using indirect rewards from metadata to overcome supervision scarcity in rare domains.
Unified framework for iterative neural network compression combining pruning, quantization, and low-rank decomposition techniques with gradual accuracy preservation.
ATLAS: framework for deploying LLM agents in autonomous trading with dynamic prompt optimization, multi-agent coordination, and late-arriving noisy rewards.
MemoryBench: benchmark for evaluating memory and continual learning capabilities in LLM systems, addressing limitations of scaling-only approaches.
Memory-augmented framework enabling LLM agents to learn classification functions from labeled examples without parameter updates using episodic memory and critiques.
Sentra-Guard: real-time defense system detecting jailbreak and prompt injection attacks on LLMs using FAISS embeddings and fine-tuned classifiers.
PORTool: algorithm for training multi-tool LLM agents using importance-weighted policy optimization to resolve credit-assignment ambiguity in tool-use decisions.
Risk-aware LLM-based agentic framework for 6G network negotiation addressing uncertainty neglect and tail-event risks.
Multi-stage computational materials discovery workflow combining generative models and graph neural networks generating 119M candidate structures.
Audit of LAION-Aesthetics Predictor examining whose cultural values and aesthetic preferences are encoded in image quality models.
ChatEHR system enabling LLM integration with electronic health records for clinical documentation automations and interactive use.
Comparative evaluation of small language models on multi-turn customer-service QA using synthetic data in resource-constrained settings.
Research showing LLMs struggle to utilize representations learned in-context, limiting few-shot adaptation capabilities.
Backdoor attack method targeting spiking neural networks via adversarial manipulation of spiking neuron hyperparameters.
Vision-language model framework for early wildfire detection and risk assessment combining satellite imagery analysis with AI.
Method using spectral heat diffusion to identify abstraction boundaries in knowledge graphs for continuous resolution control.
Model editing technique for generative recommendation systems addressing cold-start item collapse without full retraining.
Study of alignment evaluation in LLMs examining concept detection vs. behavioral routing using Chinese-origin models and ablation techniques.
Research localizing policy routing mechanisms in alignment-trained LLMs, identifying gate and amplifier heads controlling refusal behavior.
Neural improvement methods for TSP that learn search policies applying local modifications to candidate solutions rather than single-shot outputs.
Framework combining reinforcement learning with LLM-guided action spaces for optimizing drug molecules while ensuring synthesizability.
Unified framework for preference optimization revealing shared decomposition across margin-based methods to preserve chosen responses while suppressing rejected ones.
Technique using weak supervision to prevent sandbagging in LLMs, eliciting best performance without reliable verification of output quality.