SHARP:单张照片在一秒内变成可实时渲染的 3D Gaussian 场
发布时间:
SHARP:单张照片在一秒内变成可实时渲染的 3D Gaussian 场
Recommended citation: Mescheder et al., Sharp Monocular View Synthesis in Less Than a Second, ICLR 2026
Download Paper
发布时间:
Recommended citation: Mescheder et al., Sharp Monocular View Synthesis in Less Than a Second, ICLR 2026
Download Paper
发布时间:
Recommended citation: Zhang et al., Efficiently Reconstructing Dynamic Scenes One D4RT at a Time, arXiv:2512.08924, 2025.
Download Paper
发布时间:
发布时间:
发布时间:
Recommended citation: Ning et al., Code as Agent Harness: Toward Executable, Verifiable, and Stateful Agent Systems, arXiv:2605.18747, 2026.
Download Paper
发布时间:
发布时间:
Recommended citation: Ailing Zeng et al. LPM 1.0: Video-based Character Performance Model. arXiv:2604.07823v2, 2026.
Download Paper
发布时间:
FLOAT 论文精读:把 audio-driven talking portrait 的生成目标从 pixel video 换到 motion latent trajectory,再用 Flow Matching 快速采样。
Recommended citation: Ki, Min, and Chae. FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait. ICCV 2025.
Download Paper
发布时间:
发布时间:
Recommended citation: Wan Team, Alibaba Group. Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models. arXiv:2606.25041, 2026.
Download Paper
发布时间:
Recommended citation: DreamX Team et al., DreamX-World 1.0: A General-Purpose Interactive World Model, arXiv:2606.16993, 2026.
Download Paper
发布时间:
发布时间:
Recommended citation: Ailing Zeng et al. LPM 1.0: Video-based Character Performance Model. arXiv:2604.07823v2, 2026.
Download Paper
发布时间:
FLOAT 论文精读:把 audio-driven talking portrait 的生成目标从 pixel video 换到 motion latent trajectory,再用 Flow Matching 快速采样。
Recommended citation: Ki, Min, and Chae. FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait. ICCV 2025.
Download Paper
发布时间:
发布时间:
Recommended citation: DreamX Team et al., DreamX-World 1.0: A General-Purpose Interactive World Model, arXiv:2606.16993, 2026.
Download Paper
发布时间:
Recommended citation: Ning et al., Code as Agent Harness: Toward Executable, Verifiable, and Stateful Agent Systems, arXiv:2605.18747, 2026.
Download Paper
发布时间:
Recommended citation: Hirokatsu Kataoka et al., CVPR 2026 Report (Finalized ver.), LIMIT.Lab / cvpaper.challenge / Visual Geometry Group, 2026.
Download Paper
发布时间:
Recommended citation: Zhang et al., Efficiently Reconstructing Dynamic Scenes One D4RT at a Time, arXiv:2512.08924, 2025.
Download Paper
发布时间:
Recommended citation: Ma et al., StyleTalk: One-shot Talking Head Generation with Controllable Speaking Styles, AAAI 2023 Oral.
Download Paper
发布时间:
Recommended citation: Wang et al., PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation, CVPR 2026.
Download Paper
发布时间:
Recommended citation: Hirokatsu Kataoka et al., CVPR 2026 Report (Finalized ver.), LIMIT.Lab / cvpaper.challenge / Visual Geometry Group, 2026.
Download Paper
发布时间:
发布时间:
发布时间:
Recommended citation: ChatGPT shared conversation, 领域生成数学原理, retrieved 2026-06-23; supplemented with primary papers and official arXiv/OpenReview sources.
发布时间:
Recommended citation: SeFi-Team, SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion, arXiv:2606.22568, 2026.
Download Paper
发布时间:
Recommended citation: DreamX Team et al., DreamX-World 1.0: A General-Purpose Interactive World Model, arXiv:2606.16993, 2026.
Download Paper
发布时间:
Recommended citation: Lu et al., PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion, arXiv:2605.23902, 2026.
Download Paper
发布时间:
Recommended citation: Singh et al., Improved Baselines with Representation Autoencoders, arXiv:2605.18324, 2026.
Download Paper
发布时间:
Recommended citation: Li et al. Dual Diffusion for Unified Image Generation and Understanding. arXiv:2501.00289, 2024.
Download Paper
发布时间:
发布时间:
发布时间:
发布时间:
Recommended citation: Hu et al., ELF: Embedded Language Flows, arXiv:2605.10938, 2026.
Download Paper
发布时间:
Recommended citation: Pengqi Lu. Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers. arXiv:2605.06169, 2026.
Download Paper
发布时间:
Recommended citation: Ailing Zeng et al. LPM 1.0: Video-based Character Performance Model. arXiv:2604.07823v2, 2026.
Download Paper
发布时间:
发布时间:
发布时间:
发布时间:
Recommended citation: Yang et al., GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting. arXiv:2607.00959v1, 2026.
Download Paper
发布时间:
Recommended citation: Cui et al., AvatarForcing: One-Step Streaming Talking Avatars via Local-Future Sliding-Window Denoising. arXiv:2603.14331, 2026.
Download Paper
发布时间:
Recommended citation: Ma et al., StyleTalk: One-shot Talking Head Generation with Controllable Speaking Styles, AAAI 2023 Oral.
Download Paper
发布时间:
Recommended citation: Wang et al., PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation, CVPR 2026.
Download Paper
发布时间:
Recommended citation: ChatGPT shared conversation, 领域生成数学原理, retrieved 2026-06-23; supplemented with primary papers and official arXiv/OpenReview sources.
发布时间:
Recommended citation: Yang et al., GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting. arXiv:2607.00959v1, 2026.
Download Paper
发布时间:
Recommended citation: Wang et al., PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation, CVPR 2026.
Download Paper
发布时间:
Recommended citation: ChatGPT shared conversation, 领域生成数学原理, retrieved 2026-06-23; supplemented with primary papers and official arXiv/OpenReview sources.
发布时间:
Recommended citation: Black Forest Labs, FLUX.2: Analyzing and Enhancing the Latent Space of FLUX -- Representation Comparison, 2025.
Download Paper
发布时间:
Recommended citation: Singh et al., Improved Baselines with Representation Autoencoders, arXiv:2605.18324, 2026.
Download Paper
发布时间:
Recommended citation: Hu et al., ELF: Embedded Language Flows, arXiv:2605.10938, 2026.
Download Paper
发布时间:
Recommended citation: Zhengyang Geng et al. Mean Flows for One-step Generative Modeling. NeurIPS 2025.
Download Paper
发布时间:
Recommended citation: Pengqi Lu. Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers. arXiv:2605.06169, 2026.
Download Paper
发布时间:
Recommended citation: Zhen Fang et al. Flow-OPD: On-Policy Distillation for Flow Matching Models. arXiv:2605.08063, 2026.
Download Paper
发布时间:
发布时间:
Recommended citation: Pedro M. P. Curvo, Maksim Zhdanov, Floor Eijkelboom, and Jan-Willem van de Meent. Follow the Mean: Reference-Guided Flow Matching. arXiv:2605.10302, 2026.
Download Paper
发布时间:
Recommended citation: Hila Chefer, Patrick Esser, Dominik Lorenz, Dustin Podell, Vikash Raja, Vinh Tong, Antonio Torralba, Robin Rombach. Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis. arXiv:2603.06507, 2026.
Download Paper
发布时间:
发布时间:
发布时间:
FLOAT 论文精读:把 audio-driven talking portrait 的生成目标从 pixel video 换到 motion latent trajectory,再用 Flow Matching 快速采样。
Recommended citation: Ki, Min, and Chae. FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait. ICCV 2025.
Download Paper
发布时间:
Recommended citation: Black Forest Labs, FLUX.2: Analyzing and Enhancing the Latent Space of FLUX -- Representation Comparison, 2025.
Download Paper
发布时间:
Recommended citation: Yılmaz et al., Edit2Restore: Few-Shot Image Restoration via Parameter-Efficient Adaptation of Pre-trained Editing Models, arXiv 2026
Download Paper
发布时间:
Recommended citation: Pedro M. P. Curvo, Maksim Zhdanov, Floor Eijkelboom, and Jan-Willem van de Meent. Follow the Mean: Reference-Guided Flow Matching. arXiv:2605.10302, 2026.
Download Paper
发布时间:
Recommended citation: Hila Chefer, Patrick Esser, Dominik Lorenz, Dustin Podell, Vikash Raja, Vinh Tong, Antonio Torralba, Robin Rombach. Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis. arXiv:2603.06507, 2026.
Download Paper
发布时间:
Recommended citation: Yang et al., GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting. arXiv:2607.00959v1, 2026.
Download Paper
发布时间:
Recommended citation: Zhengyang Geng et al. Mean Flows for One-step Generative Modeling. NeurIPS 2025.
Download Paper
发布时间:
发布时间:
Recommended citation: Hirokatsu Kataoka et al., CVPR 2026 Report (Finalized ver.), LIMIT.Lab / cvpaper.challenge / Visual Geometry Group, 2026.
Download Paper
发布时间:
Recommended citation: Zhengyang Geng et al. Mean Flows for One-step Generative Modeling. NeurIPS 2025.
Download Paper
发布时间:
Recommended citation: Pedro M. P. Curvo, Maksim Zhdanov, Floor Eijkelboom, and Jan-Willem van de Meent. Follow the Mean: Reference-Guided Flow Matching. arXiv:2605.10302, 2026.
Download Paper
发布时间:
Recommended citation: Hila Chefer, Patrick Esser, Dominik Lorenz, Dustin Podell, Vikash Raja, Vinh Tong, Antonio Torralba, Robin Rombach. Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis. arXiv:2603.06507, 2026.
Download Paper
发布时间:
Recommended citation: ChatGPT shared conversation, 领域生成数学原理, retrieved 2026-06-23; supplemented with primary papers and official arXiv/OpenReview sources.
发布时间:
Recommended citation: Yılmaz et al., Edit2Restore: Few-Shot Image Restoration via Parameter-Efficient Adaptation of Pre-trained Editing Models, arXiv 2026
Download Paper
发布时间:
发布时间:
Recommended citation: SeFi-Team, SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion, arXiv:2606.22568, 2026.
Download Paper
发布时间:
Recommended citation: Black Forest Labs, FLUX.2: Analyzing and Enhancing the Latent Space of FLUX -- Representation Comparison, 2025.
Download Paper
发布时间:
Recommended citation: Lu et al., PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion, arXiv:2605.23902, 2026.
Download Paper
发布时间:
Recommended citation: Singh et al., Improved Baselines with Representation Autoencoders, arXiv:2605.18324, 2026.
Download Paper
发布时间:
Recommended citation: Yılmaz et al., Edit2Restore: Few-Shot Image Restoration via Parameter-Efficient Adaptation of Pre-trained Editing Models, arXiv 2026
Download Paper
发布时间:
发布时间:
发布时间:
Recommended citation: Hu et al., ELF: Embedded Language Flows, arXiv:2605.10938, 2026.
Download Paper
发布时间:
Recommended citation: Lu et al., PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion, arXiv:2605.23902, 2026.
Download Paper
发布时间:
Recommended citation: SeFi-Team, SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion, arXiv:2606.22568, 2026.
Download Paper
发布时间:
Recommended citation: Black Forest Labs, FLUX.2: Analyzing and Enhancing the Latent Space of FLUX -- Representation Comparison, 2025.
Download Paper
发布时间:
Recommended citation: Ning et al., Code as Agent Harness: Toward Executable, Verifiable, and Stateful Agent Systems, arXiv:2605.18747, 2026.
Download Paper
发布时间:
RoPE 外推友好性完整解析
发布时间:
发布时间:
发布时间:
发布时间:
发布时间:
发布时间:
发布时间:
发布时间:
发布时间:
发布时间:
发布时间:
发布时间:
发布时间:
发布时间:
发布时间:
Recommended citation: Pengqi Lu. Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers. arXiv:2605.06169, 2026.
Download Paper
发布时间:
Recommended citation: Wan Team, Alibaba Group. Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models. arXiv:2606.25041, 2026.
Download Paper
发布时间:
Recommended citation: Li et al. Dual Diffusion for Unified Image Generation and Understanding. arXiv:2501.00289, 2024.
Download Paper
发布时间:
发布时间:
Recommended citation: Zhen Fang et al. Flow-OPD: On-Policy Distillation for Flow Matching Models. arXiv:2605.08063, 2026.
Download Paper
发布时间:
Recommended citation: Yang et al., GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting. arXiv:2607.00959v1, 2026.
Download Paper
发布时间:
Recommended citation: Cui et al., AvatarForcing: One-Step Streaming Talking Avatars via Local-Future Sliding-Window Denoising. arXiv:2603.14331, 2026.
Download Paper
发布时间:
FLOAT 论文精读:把 audio-driven talking portrait 的生成目标从 pixel video 换到 motion latent trajectory,再用 Flow Matching 快速采样。
Recommended citation: Ki, Min, and Chae. FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait. ICCV 2025.
Download Paper
发布时间:
Recommended citation: Hu et al., ELF: Embedded Language Flows, arXiv:2605.10938, 2026.
Download Paper
发布时间:
发布时间:
Recommended citation: Zhang et al., Efficiently Reconstructing Dynamic Scenes One D4RT at a Time, arXiv:2512.08924, 2025.
Download Paper
发布时间:
Recommended citation: Pedro M. P. Curvo, Maksim Zhdanov, Floor Eijkelboom, and Jan-Willem van de Meent. Follow the Mean: Reference-Guided Flow Matching. arXiv:2605.10302, 2026.
Download Paper
发布时间:
Recommended citation: ChatGPT shared conversation, 领域生成数学原理, retrieved 2026-06-23; supplemented with primary papers and official arXiv/OpenReview sources.
发布时间:
Recommended citation: Zhen Fang et al. Flow-OPD: On-Policy Distillation for Flow Matching Models. arXiv:2605.08063, 2026.
Download Paper
发布时间:
Recommended citation: Black Forest Labs, FLUX.2: Analyzing and Enhancing the Latent Space of FLUX -- Representation Comparison, 2025.
Download Paper
发布时间:
Recommended citation: Singh et al., Improved Baselines with Representation Autoencoders, arXiv:2605.18324, 2026.
Download Paper
发布时间:
Recommended citation: Hila Chefer, Patrick Esser, Dominik Lorenz, Dustin Podell, Vikash Raja, Vinh Tong, Antonio Torralba, Robin Rombach. Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis. arXiv:2603.06507, 2026.
Download Paper
发布时间:
发布时间:
Recommended citation: Hirokatsu Kataoka et al., CVPR 2026 Report (Finalized ver.), LIMIT.Lab / cvpaper.challenge / Visual Geometry Group, 2026.
Download Paper
发布时间:
Recommended citation: SeFi-Team, SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion, arXiv:2606.22568, 2026.
Download Paper
发布时间:
Recommended citation: Ning et al., Code as Agent Harness: Toward Executable, Verifiable, and Stateful Agent Systems, arXiv:2605.18747, 2026.
Download Paper
发布时间:
Recommended citation: Wan Team, Alibaba Group. Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models. arXiv:2606.25041, 2026.
Download Paper
发布时间:
Recommended citation: Cui et al., AvatarForcing: One-Step Streaming Talking Avatars via Local-Future Sliding-Window Denoising. arXiv:2603.14331, 2026.
Download Paper
发布时间:
Recommended citation: Ma et al., StyleTalk: One-shot Talking Head Generation with Controllable Speaking Styles, AAAI 2023 Oral.
Download Paper
发布时间:
Recommended citation: Lu et al., PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion, arXiv:2605.23902, 2026.
Download Paper
发布时间:
发布时间:
发布时间:
Recommended citation: Cui et al., AvatarForcing: One-Step Streaming Talking Avatars via Local-Future Sliding-Window Denoising. arXiv:2603.14331, 2026.
Download Paper
发布时间:
Recommended citation: Yang et al., GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting. arXiv:2607.00959v1, 2026.
Download Paper
发布时间:
Recommended citation: Ma et al., StyleTalk: One-shot Talking Head Generation with Controllable Speaking Styles, AAAI 2023 Oral.
Download Paper
发布时间:
Recommended citation: Wang et al., PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation, CVPR 2026.
Download Paper
发布时间:
FLOAT 论文精读:把 audio-driven talking portrait 的生成目标从 pixel video 换到 motion latent trajectory,再用 Flow Matching 快速采样。
Recommended citation: Ki, Min, and Chae. FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait. ICCV 2025.
Download Paper
发布时间:
Recommended citation: SeFi-Team, SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion, arXiv:2606.22568, 2026.
Download Paper
发布时间:
Recommended citation: Li et al. Dual Diffusion for Unified Image Generation and Understanding. arXiv:2501.00289, 2024.
Download Paper
发布时间:
Recommended citation: Zhen Fang et al. Flow-OPD: On-Policy Distillation for Flow Matching Models. arXiv:2605.08063, 2026.
Download Paper
发布时间:
Recommended citation: Pengqi Lu. Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers. arXiv:2605.06169, 2026.
Download Paper
发布时间:
Recommended citation: Cui et al., AvatarForcing: One-Step Streaming Talking Avatars via Local-Future Sliding-Window Denoising. arXiv:2603.14331, 2026.
Download Paper
发布时间:
Recommended citation: Wan Team, Alibaba Group. Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models. arXiv:2606.25041, 2026.
Download Paper
发布时间:
Recommended citation: ChatGPT shared conversation, 领域生成数学原理, retrieved 2026-06-23; supplemented with primary papers and official arXiv/OpenReview sources.
发布时间:
Recommended citation: DreamX Team et al., DreamX-World 1.0: A General-Purpose Interactive World Model, arXiv:2606.16993, 2026.
Download Paper
发布时间:
发布时间:
Recommended citation: Ailing Zeng et al. LPM 1.0: Video-based Character Performance Model. arXiv:2604.07823v2, 2026.
Download Paper
发布时间:
发布时间:
发布时间:
Recommended citation: Zhang et al., Efficiently Reconstructing Dynamic Scenes One D4RT at a Time, arXiv:2512.08924, 2025.
Download Paper
发布时间:
Recommended citation: Mescheder et al., Sharp Monocular View Synthesis in Less Than a Second, ICLR 2026
Download Paper
发布时间:
Recommended citation: Li et al. Dual Diffusion for Unified Image Generation and Understanding. arXiv:2501.00289, 2024.
Download Paper
发布时间:
Recommended citation: Hirokatsu Kataoka et al., CVPR 2026 Report (Finalized ver.), LIMIT.Lab / cvpaper.challenge / Visual Geometry Group, 2026.
Download Paper
发布时间:
Recommended citation: Wan Team, Alibaba Group. Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models. arXiv:2606.25041, 2026.
Download Paper
发布时间:
Recommended citation: DreamX Team et al., DreamX-World 1.0: A General-Purpose Interactive World Model, arXiv:2606.16993, 2026.
Download Paper
发布时间: