CultureVLM characterizes and improves cultural understanding of vision-language models across more than 100 countries using culturally-grounded benchmarks and training procedures.
MASCOT is a multi-agent framework for socio-collaborative companions that uses bi-level optimization—persona-aware behavioral alignment and collaborative dialogue optimization—to counter persona collapse and social sycophancy, improving role consistency and reducing redundant dialogue.
A systematic study showing that reasoning capabilities alone are insufficient for LLMs in multi-turn mental health conversations, isolating failure modes that demand additional safety and empathy-aware design.
SARA is a unified RAG framework that balances local factual precision with global coverage by combining natural-language spans with compact semantic compression vectors, achieving consistent gains under strict context budgets.
MedHalu is a fine-grained benchmark for studying hallucinations in LLM responses to consumer healthcare queries, analyzing hallucination patterns across models, query types, and medical specialties.
A unified study of LLM self-distillation that combines teacher agreement, EMA stabilization, contrastive learning, feature matching, and divergence clipping to improve adaptation without stronger external teachers.
Sharpness-aware prompt evolution optimizes prompts for both performance and robustness by penalizing sharp regions in the prompt loss landscape, yielding prompts that transfer better across tasks and LLM families.
Sysformer learns adaptive, query-conditioned system prompts to safeguard frozen large language models, providing fine-grained safety control without modifying model weights.
Distills multi-agent debate into one LLM through reasoning-enhanced fine-tuning, trajectory-based augmentation, and process-aware distillation, moving computation from inference to training.
A comprehensive survey of Protein Large Language Models, organizing the field along architectures, training objectives, datasets, and downstream tasks across biology, chemistry, and medicine.
DyGET converts temporal graph embeddings into interpretable trajectories that empower interdisciplinary insights through cross-disciplinary discovery and longitudinal community analysis.
A survey deconstructing the ethics of large language models — from long-standing issues such as bias and misinformation to newly emerging dilemmas around agentic behavior and cross-cultural alignment.
A survey of efficient LLM training organized around data-centric techniques — selection, mixing, ordering, and synthesis — and their trade-offs with compute and downstream performance.
ProteinGPT is a multimodal LLM that integrates protein sequence and structural representations in a unified generative interface for property prediction and structure understanding.
We propose a framework for developing topology-aware Multi-Agent Systems (MAS), emphasizing agent selection, structure profiling, and topology synthesis, to enhance coordination and efficiency in complex task.
ScreenLLM introduces a stateful screen schema and key-frame extractor that compresses dynamic UI sessions into time-aware summaries, enabling efficient GUI understanding and action prediction with multimodal LLMs.
Studies how rewriting YouTube video titles when sharing them on Reddit affects engagement, using controlled comparisons to isolate the effects of language and community context.
A methodology-centered survey of LLM agents, connecting agent architectures, collaboration, and evolution with evaluation, tools, applications, and open challenges.
UniGuard is a universal safety guardrail for multimodal LLMs, defending against cross-modal jailbreak attacks across image and text channels with low utility cost.
RNA-GPT is a multimodal generative system that combines RNA sequence reasoning with structural cues for property prediction, retrieval, and natural-language interaction over RNA data.
Peer review is fundamental to the integrity and advancement of scientific publication. Traditional methods of peer review analyses often rely on exploration and statistics of existing peer review data...
PrivacyMind teaches LLMs to be contextual privacy protection learners that recognize sensitive content in context and adapt outputs accordingly, preserving utility while reducing leakage.
We address fairness issues in graph anomaly detection, providing benchmark datasets and comprehensive evaluation frameworks for fair anomaly detection on graphs....
We present a framework and benchmark to evaluate LLMs' multilingual capabilities in healthcare queries, revealing significant performance gaps across languages and providing insights for improving hea...
Social media platforms are hubs for multimodal information exchange, encompassing text, images, and videos, making it challenging for machines to comprehensively understand the information. Multimodal...
We propose a semi-offline reinforcement learning approach for optimizing text generation in language models, balancing exploration and exploitation effectively....
CODER is a graph-based code recommendation framework for open source developers, modeling heterogeneous OSS contribution networks to deliver multi-modal recommendations across realistic workloads.
We develop methods to predict how information spreads across different online communities, revealing patterns in cross-platform information diffusion....
We propose prototypical fine-tuning, a novel framework for fine-tuning pretrained language models that maintains robust performance across varying data sizes.
A novel subgraph reasoning paradigm for fake news detection that provides explainability while improving generalization through reinforcement learning and hierarchical graph attention networks.
We propose a fine-grained reasoning framework for fake news detection by following the human information-processing model and designing a prior-aware bi-channel kernel graph network....