what's new in generative ai
recent papers in generative ai, each with a practical, plain-language summary. models that create.
want the foundations first?take the generative ai learning path →
- 📄 paperJul 2026
Set Diffusion: Interpolating Token Orderings Between Autoregression and Diffusion for Fast and Flexible Decoding
Marianne Arriola, Volodymyr Kuleshov
this paper presents set diffusion, a method that combines autoregressive and diffusion models for faster and more flexible decoding of discrete tokens. ml engineers working with sequence generation can use this to improve the efficiency and control over their generative models.
- 📄 paperJul 2026
A Physics-guided Fine-tuned LLM-based Framework for Customized Power Distribution System Feeder Generation
Zhenghao Zhou, Yiyan Li, Tao Xu +3
this research proposes a physics-guided llm framework to generate customized power distribution system feeders. this can help engineers design more efficient and reliable power grids by automating the generation of complex system configurations.
- 📄 paperJun 2026
Unified Multimodal Autoregressive Modeling with Shared Context—Visual Tokenizer is Key to Unification
Wujian Peng, Lingchen Meng, Yuxuan
this paper presents a unified multimodal autoregressive model that uses a shared context and visual tokenizer for unification. ai developers can use this approach to build more cohesive and capable multimodal models that can process and generate across different data types like text and images.
- 📄 paperJun 2026
Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards
Ritesh Thawkar, Shravan Venkatraman, Omkar Thawakar +5
this paper proposes a self-evolving framework for unified multimodal understanding and generation using self-consistency rewards. this approach allows models to improve their ability to both comprehend and create multimodal content, leading to more robust and versatile ai systems.
- 📄 paperJun 2026
Flex4DHuman: Flexible Multi-view Video Diffusion for 4D Human Reconstruction
Jen-Hao Cheng, Yipeng Wang, Hao Zhang +2
this paper presents flex4dhuman, a multi-view video diffusion model for flexible 4d human reconstruction. for practitioners in computer graphics, animation, or virtual try-on, this provides a powerful tool for generating realistic and dynamic 3d human models from video input.
- 📄 paperJun 2026
TooBad: Backdoor Diffusion Models with Ultra-Low Poison Rate and Imperceptible Trigger
Vu Tuan Truong, Long Bao Le
this paper reveals a new vulnerability, "toobad," demonstrating how diffusion models can be backdoored with imperceptible triggers and low poison rates. for ai security researchers and developers, this highlights critical security risks in generative models, emphasizing the need for robust defenses against malicious attacks.
- 📄 paperJun 2026
Generative artificial intelligence creates delicious, sustainable, and nutritious burgers
Jingxuan Li, Yuxuan Liang, Yifan Li +4
this paper shows how generative ai can learn human taste preferences from recipe data to design foods that are delicious, nutritious, and sustainable. for practitioners, this means ai can be used to innovate in food product development, potentially reducing development cycles and improving health and environmental outcomes.
- 📄 paperJun 2026
Accelerating Speculative Diffusions via Block Verification
Alexander Soen, Hisham Husain, Valentin De Bortoli +1
this paper presents a method to accelerate speculative decoding for large language models using block verification. for practitioners, this means faster inference times for llms, which can significantly reduce computational costs and improve user experience in production environments.
- 📄 paperJun 2026
A2D2: Fine-Tuning Any-Length Discrete Diffusion for Adaptive Decoding
Sophia Tang, Yuchen Zhu, Molei Tao +1
this paper presents a2d2, a method for fine-tuning any-length discrete diffusion models for adaptive decoding in sequence generation. for practitioners in nlp or sequence modeling, this offers a stable and flexible framework for generating sequences of varying lengths, improving the adaptability of diffusion models.
- 📄 paperJun 2026
MAP: evaluation and multi-agent enhancement of large language models for inpatient pathways
evaluates and improves LLMs for complex inpatient decision-making by moving beyond question-answering to real clinical workflows. practitioners need this because it addresses the gap between medical benchmarks and the messy, multi-step reasoning required in actual hospital settings.
- 📄 paperMay 2026
Token Time Continuous Diffusion for Language Modeling
Parikshit Bansal, Sujay Sanghavi
this paper introduces token time continuous diffusion (ttcd), a novel diffusion language model operating in continuous space. researchers in nlp can explore this method as an alternative to traditional autoregressive models for generating text, potentially offering new ways to control and synthesize language.
- 📄 paperMay 2026
Qwen-Image-2.0 Technical Report
Qwen Team
this technical report introduces qwen-image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise editing. for practitioners in computer vision and creative fields, this model offers advanced capabilities for generating and manipulating images within a single framework, addressing challenges like ultra-long text prompts.