Welcome to UNIST<br>International Affairs Team

Welcome to UNIST
International Affairs Team

/

News

Research

Understanding Why Multimodal AI Learns More Robustly

Abstract A surge of recent advancements has consistently highlighted the superiority of multimodal learning over unimodal approaches across a variety of tasks. However, the theoretical foundations elucidating this advantage remain underexplored: existing theoretical analyzes are often constrained by tight assumptions, and lack empirical validation. In this paper, we bridge this gap by proposing a novel theoretical framework grounded in convolutional smoothing, offering a new perspective on how multimodal learning contributes to a smoother loss landscape compared to unimodal learning. Building upon this theoretical foundation, we introduce a simple yet effective distributional training strategy based on stochastic modality pairing instead of a fixed pairing; Thus, further promoting a flatter landscape via convolutional smoothing. Our empirical results across various multimodal datasets demonstrate that multimodal models not only achieve higher performance but also exhibit flatter loss landscape, which represent better generalization and robustness. Artificial intelligence models trained on combinations of images, speech, and text consistently outperform those that learn from a single type of data. While this advantage has been widely observed, the reason multimodal learning produces more accurate and reliable models has remained largely unexplained. Researchers at UNIST have now provided a mathematical explanation for this phenomenon and used it to develop a new training strategy that further improves multimodal learning. The study, led by Professor Sung Whan Yoon of the Graduate School of Artificial Intelligence, shows that learning from multiple data modalities produces a flatter loss landscape—a property associated with models that generalize more effectively to unfamiliar data and remain more resilient to noise and perturbations. The researchers demonstrate that this effect arises through a mathematical mechanism known as convolutional smoothing. By learning from different types of information simultaneously, multimodal models naturally smooth abrupt changes in the optimization process, creating a more stable learning landscape than models trained on a single modality. Building on this theoretical framework, the team developed Distributional Multimodal Learning (DML), a simple training strategy that replaces fixed image-text or image-audio pairs with randomly sampled combinations drawn from the same class. This broader variety of training examples further enhances the smoothing effect, leading to stronger generalization and more robust learning. Across multiple multimodal benchmark datasets, DML consistently outperformed conventional multimodal training. The method improved both classification accuracy and cross-modal retrieval tasks, including matching images with their corresponding text descriptions and retrieving images from textual queries. “Our work provides a theoretical foundation for understanding why multimodal learning consistently outperforms unimodal learning,” the research team said. “It also demonstrates how that understanding can be translated into a simple yet effective training strategy that improves both robustness and generalization.” The study's first author is Jaejun Lee, a researcher in the UNIST Graduate School of Artificial Intelligence. The work has been accepted for presentation at the International Conference on Machine Learning (ICML) 2026, one of the world's leading conferences on artificial intelligence, to be held in Seoul from July 6–11. The research was supported by the National Research Foundation of Korea (NRF) and the Institute for Information & Communications Technology Planning &Evaluation (IITP) through programs funded by the Ministry of Science and ICT. Journal Reference Jae-Jun Lee and Sung Whan Yoon, "Understanding Multimodal Learning: A Loss Landscape Smoothness Perspective," ICML '26 ., (2026).

Understanding Why Multimodal AI Learns More Robustly

Research

New Deep Learning Model Extends Reliable Wildfire Forecasts Beyond Two Weeks

Abstract The 2025 Los Angeles wildfires highlighted growing risks from climate-driven extremes and the need for reliable wildfire forecasting beyond short lead times. Although the European Center for Medium-Range Weather Forecasts provides global fire weather index forecasts, these products are primarily designed for long-range climate outlooks, and their practical effectiveness remains uncertain in regions with limited local forecasting infrastructure. Here we show that a global deep learning framework, forecasting daily fire weather index values up to 31 days ahead, consistently improves forecast accuracy and reduces bias relative to operational numerical forecasts. The framework learns nonlinear and lagged fire–weather relationships by integrating past fire weather index dynamics and future meteorological conditions. Importantly, forecast bias is reduced in 85% of grid cells where high wildfire exposure coincides with high socioeconomic vulnerability. This demonstrates that data-driven forecasting can help bridge critical information gaps in underserved regions and support more equitable climate risk management. Reliable wildfire forecasts become increasingly difficult beyond the first few weeks, limiting their value for medium-range decision-making. A research team, led by Professor Jungho of the Department of Civil, Urban, Earth, and Environmental Engineering at UNIST, has shown that a global deep learning framework can extend reliable daily wildfire forecasts to 31 days while reducing forecast bias compared with existing operational forecasting systems. Known as FWI-Net, the framework predicts daily values of the Fire Weather Index (FWI), a widely used measure of wildfire danger. Unlike conventional forecasting methods, it combines historical fire-weather conditions with future meteorological forecasts, allowing it to capture the cumulative effects of prolonged heat and drought on future wildfire risk. Across the full 31-day forecasting period, FWI-Net reduced prediction error by 6.6% compared with the operational forecasting system of the European Center for Medium-Range Weather Forecasts (ECMWF). During the first week, prediction error decreased by 12.4%. The model also extended the period of meaningful forecasts under very high wildfire danger by five days. The model performed particularly well in regions where high wildfire exposure coincides with high socioeconomic vulnerability. Forecast bias was reduced across 85% of these areas, and in regions with limited forecasting infrastructure, FWI-Net maintained useful forecasting skill for an average of 22 days. “By combining historical fire-weather conditions with future weather forecasts, the model captures patterns that conventional forecasting systems often miss,” the research team said. “Most importantly, it provides more reliable forecasts for regions facing high wildfire risk despite limited forecasting infrastructure.” “As climate change increases wildfire risk around the world, reliable forecasting is becoming essential for disaster preparedness,” said Professor Im. “We expect this framework to support medium-range wildfire planning while helping reduce information gaps in regions with limited forecasting capacity.” The study was co-led by Professor Yoojin Kang of Kookmin University and Sihyun Lee of UNIST, who served as first authors. The findings were published online in Communications Earth & Environment on June 28, 2026. The research was supported by the Ministry of Environment (ME), the Korea Forest Service, and the National Research Foundation of Korea (NRF). Journal Reference Yoojin Kang, Sihyun Lee, Dongjin Cho, and Jungho Im, "Deep learning-based forecasting provides a pathway to closing wildfire information gaps in underserved regions," Commun. Earth Environment. , (2026).

New Deep Learning Model Extends Reliable Wildfire Forecasts Beyond Two Weeks

News

UNIST Unveils Integrated Quantum Ecosystem at Quantum Korea 2026

UNIST presented its vision for the next phase of quantum research and education at Quantum Korea 2026 , introducing an integrated ecosystem for quantum innovation. The initiative brings together the newly established Graduate School of Quantum Science and Technology, the Institute of UNIST Quantum Information ( Tentative Name )—known as Uni<Q>, and the Quantum-Nano FAB. Together they support the full spectrum of quantum research at UNIST—from educating future specialists and advancing fundamental science to enabling quantum device fabrication and collaboration with industry. At its Uni<Q> exhibition booth during the Quantum Korea 2026 , held July 2–4 at Dongdaemun Design Plaza (DDP) in Seoul, UNIST highlighted research in quantum computing, sensing, advanced materials, and devices while introducing the university's integrated approach to quantum education and research. More than 1,000 researchers, industry representatives, students, and members of the public visited the UNIST booth during the three-day event. Faculty members and graduate students also meet with prospective applicants to introduce the new curriculum, research opportunities, and career pathways. The event also highlighted UNIST's growing research leadership in quantum science. During the opening ceremony, Professor Je-Hyung Kim of the Graduate School of Quantum Science and Technology received a Commendation from the Deputy Prime Minister and Minister of Science and ICT for his contributions to quantum science and technology. “Progress in quantum technology requires research infrastructure, talented people, and sustained work on fundamental challenges to advance together,” said Ki-Bok Park, Dean of the Graduate School of Quantum Science and Technology at UNIST. “Our goal is to educate researchers who understand the principles of quantum science and can apply that knowledge to practical technological challenges.” Professor Kim added, “This recognition reflects the research capabilities UNIST has built in quantum science and technology, as well as the potential for further growth. We will continue advancing foundational technologies and expanding the impact of our research.” Held under the theme “Quantum Becomes Reality: Bold Challenges for Innovation,” Quantum Korea 2026 brought together 56 companies and research organizations from 12 countries to share developments in quantum computing, communication, sensing, and commercialization.

UNIST Unveils Integrated Quantum Ecosystem at Quantum Korea 2026

Research

New Training Method Improves Knowledge Distillation for Compact Generative AI

Abstract We demonstrate that in knowledge distillation for diffusion models, the teacher network's highly complex denoising process—stemming from its substantially larger capacity—poses a significant challenge for the student model to faithfully mimic. To address this problem, we propose a coarse-to-fine distillation framework with LInear Fitting-based distillation (LIFT) and Piecewise Local Adaptive Coefficient Estimation (PLACE). First, LIFT decomposes the objective into a coarse'' alignment and afine'' refinement. The student is then trained on coarse alignment before proceeding to hard refinement. Second, LInear Fitting-based distillation extends LIFT to address spatially non-uniform errors by partitioning outputs into error-based groups, providing locally adaptive guidance. Our comprehensive experimental results demonstrate that ours, \core~with \pick, outperforms previous knowledge distillation on diffusion models based on both U-Net and DiT architectures. Furthermore, as compression rates become exceedingly high, conventional knowledge distillation fails to provide sufficient guidance, thereby preventing lightweight diffusion models from achieving stable training. In contrast, our method demonstrates stable convergence even under such extreme compression ratios. Compact generative AI models promise to bring image generation to personal devices, but reducing model size often comes at the expense of image quality and training stability. Researchers at UNIST have developed a new knowledge distillation framework that enables lightweight diffusion models to learn more effectively from much larger models without increasing inference costs. Led by Professor Jaejun Yoo of the Graduate School of Artificial Intelligence, the study addresses a key challenge in knowledge distillation—the process of transferring the capabilities of a large AI model to a smaller one. The researchers found that as teacher models become increasingly capable, the complexity of their denoising process makes it progressively more difficult for compact student models to reproduce faithfully. Rather than asking the student model to reproduce every detail from the outset, the framework restructures learning into successive stages, allowing it to first capture an image's overall structure before progressively refining finer details. The approach further improves learning by providing additional guidance where prediction errors are greatest while adapting supervision throughout training. This strategy is implemented through two complementary techniques, LIFT (LInear Fitting-based Distillation) and PLACE (Piecewise Local Adaptive Coefficient Estimation). The approach changes only the training strategy, requiring neither additional model parameters nor extra computation during image generation. Across multiple image-generation benchmarks, the proposed framework consistently outperformed existing knowledge distillation methods. It proved effective for text-to-image synthesis, class-conditional generation, and unconditional image generation, while remaining compatible with multiple diffusion architectures, including U-Net, DiT, and MMDiT, the architecture used in Stable Diffusion 3. The advantages became particularly evident under extreme model compression, where the widening gap between teacher and student models often leaves conventional knowledge distillation unable to provide sufficient guidance for stable learning. When compressing a teacher model containing 78.7 million parameters into a student model with just 1.3 million parameters, the proposed method maintained stable convergence and generated substantially higher-quality images. “Rather than modifying model architectures, our approach improves how compact models learn from larger ones,” said Professor Yoo. “Because it introduces no additional model parameters or inference-time computation, it can be readily applied across a wide range of generative AI models, making compact generative AI practical across a wide range of computing environments.” Hyunsoo Han and Sangyeob Yeo of the Graduate School of Artificial Intelligence at UNIST contributed to the study. The findings have been accepted for presentation at the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2026. Journal Reference Hyunsoo Han, Sangyeop Yeo, and Jaejun Yoo, "LIFT and PLACE: A Simple, Stable, and Effective Knowledge Distillation Framework for Lightweight Diffusion Models," CVPR'26

New Training Method Improves Knowledge Distillation for Compact Generative AI

News

Seven Papers from UNIST Accepted to ACM DIS 2026

UNIST researchers earned international recognition through seven papers accepted to ACM Designing Interactive Systems (DIS) 2026, one of the world's leading conferences in human-computer interaction (HCI), interactive systems, and user experience (UX) design. Their research explores how emerging technologies—including AI, digital communication, and interactive media—can be designed to better support everyday life, strengthen human relationships, and encourage self-reflection. Held at the National University of Singapore, ACM DIS brings together researchers from design, computer science, and related disciplines to explore the future of human-centered technologies. Faculty members from the Department of Design at UNIST contributed research spanning embodied interaction, human-centered AI, and algorithmic transparency, demonstrating how thoughtful design can make emerging technologies more intuitive, meaningful, and accessible. Rethinking Everyday Interaction Professor Young-Woo Park's research group presented four papers examining how physical interaction can reshape familiar digital experiences—from notifications and workplace communication to messaging and music sharing. Together, the four projects demonstrate how subtle changes in interaction design can make everyday technologies more intuitive, expressive, and emotionally engaging. Prepy replaces conventional notification sounds with gentle vertical movement to communicate upcoming events, while SOTATE visualizes a person's social battery to support communication in open-plan offices. Adilet adds a tangible dimension to digital messaging, and Reelee transforms music sharing into an interactive experience that helps people maintain emotional connections across distance. Human-Centered AI Two papers from Professor Kyungho Lee's group examined how AI can support reflection, creativity, and civic participation while keeping people at the center of decision-making. Augmentiary uses large language models (LLMs) to support reflective journaling by providing thoughtful feedback that encourages self-reflection while preserving users' agency. Another study presents an AI-assisted platform for local journalism that enables citizens to contribute community news through AI-supported writing combined with collaborative human review, strengthening trust and civic engagement. Is This the Real Me? Professor Dajung Kim's research group explored how recommendation algorithms shape digital identity through TubeLens . The project asks a simple but thought-provoking question: " Is This the Real Me?" By transforming a user's YouTube recommendation history into an algorithmic self-portrait, it encourages users to critically examine how recommendation algorithms shape—and sometimes distort—their digital identities. Although the seven projects span diverse applications, they share a common goal: designing technologies that are more human-centered, transparent, and meaningful. Rather than adding complexity, the research demonstrates how thoughtful interaction design can help people engage with AI and digital systems in more intuitive and empowering ways. Together, the seven papers reflect the growing international recognition of UNIST's Department of Design in HCI and interactive design research, while highlighting the expanding role of human-centered design in shaping the future of emerging technologies.

Seven Papers from UNIST Accepted to ACM DIS 2026
UNIQUE UNISTAR