Sitemap
A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.
Pages
Jinyang Zhang
Personal website for projects, publications, foundations, blogs, and CV.
Posts
GaussianEmoTalker 深读:用 3D Gaussian 做实时情绪说话头像
发布时间:
Recommended citation: Yang et al., GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian Splatting. arXiv:2607.00959v1, 2026.
Download Paper
AvatarForcing 精读:一步流式 diffusion 如何稳住分钟级 talking avatar
发布时间:
Recommended citation: Cui et al., AvatarForcing: One-Step Streaming Talking Avatars via Local-Future Sliding-Window Denoising. arXiv:2603.14331, 2026.
Download Paper
Wan-Streamer 深读:端到端实时音视频全双工模型到底解决了什么
发布时间:
Wan-Streamer 深读:端到端实时音视频全双工模型到底解决了什么
Recommended citation: Wan Team, Alibaba Group. Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models. arXiv:2606.25041, 2026.
Download Paper
生图 / 生视频 RL 后训练:从 DPO、GRPO 到 Diffusion / Flow Alignment
发布时间:
生图 / 生视频 RL 后训练:从 DPO、GRPO 到 Diffusion / Flow Alignment
Recommended citation: ChatGPT shared conversation, 领域生成数学原理, retrieved 2026-06-23; supplemented with primary papers and official arXiv/OpenReview sources.
SeFi-Image 深读:Semantic-First Diffusion 如何把语义先行带进文生图基础模型
发布时间:
SeFi-Image 深读:Semantic-First Diffusion 如何把语义先行带进文生图基础模型
Recommended citation: SeFi-Team, SeFi-Image: A Text-to-Image Foundation Model with Semantic-First Diffusion, arXiv:2606.22568, 2026.
Download Paper
DreamX-World 1.0 深读:交互式世界模型不是视频生成,而是全栈系统工程
发布时间:
DreamX-World 1.0 深读:交互式世界模型不是视频生成,而是全栈系统工程
Recommended citation: DreamX Team et al., DreamX-World 1.0: A General-Purpose Interactive World Model, arXiv:2606.16993, 2026.
Download Paper
CVPR 2026 Report 深读:视觉研究正在从模型能力转向系统边界
发布时间:
CVPR 2026 Report 深读:视觉研究正在从模型能力转向系统边界
Recommended citation: Hirokatsu Kataoka et al., CVPR 2026 Report (Finalized ver.), LIMIT.Lab / cvpaper.challenge / Visual Geometry Group, 2026.
Download Paper
FLUX.2 表征比较:VAE latent 不是预处理,而是生成模型的接口
发布时间:
FLUX.2 表征比较:VAE latent 不是预处理,而是生成模型的接口
Recommended citation: Black Forest Labs, FLUX.2: Analyzing and Enhancing the Latent Space of FLUX -- Representation Comparison, 2025.
Download Paper
D4RT:把动态 4D 重建改写成时空点查询
发布时间:
D4RT:把动态 4D 重建改写成时空点查询
Recommended citation: Zhang et al., Efficiently Reconstructing Dynamic Scenes One D4RT at a Time, arXiv:2512.08924, 2025.
Download Paper
StyleTalk:从参考视频里抽取 speaking style 的 one-shot talking head
发布时间:
StyleTalk:从参考视频里抽取 speaking style 的 one-shot talking head
Recommended citation: Ma et al., StyleTalk: One-shot Talking Head Generation with Controllable Speaking Styles, AAAI 2023 Oral.
Download Paper
PC-Talk:用隐式关键点做可精细控制的 talking face
发布时间:
PC-Talk:用隐式关键点做可精细控制的 talking face
Recommended citation: Wang et al., PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation, CVPR 2026.
Download Paper
PiD:把 latent decoder 改成 Pixel Diffusion
发布时间:
PiD:把 latent decoder 改成 Pixel Diffusion
Recommended citation: Lu et al., PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion, arXiv:2605.23902, 2026.
Download Paper
RAEv2:Representation Autoencoder 的三个关键改进
发布时间:
RAEv2:Representation Autoencoder 的三个关键改进
Recommended citation: Singh et al., Improved Baselines with Representation Autoencoders, arXiv:2605.18324, 2026.
Download Paper
ELF:把扩散语言模型留在连续 embedding 空间里
发布时间:
ELF:把扩散语言模型留在连续 embedding 空间里
Recommended citation: Hu et al., ELF: Embedded Language Flows, arXiv:2605.10938, 2026.
Download Paper
Dual Diffusion:用扩散模型同时做图像生成和视觉理解
发布时间:
Dual Diffusion:用扩散模型同时做图像生成和视觉理解
Recommended citation: Li et al. Dual Diffusion for Unified Image Generation and Understanding. arXiv:2501.00289, 2024.
Download Paper
Seedance 2.0:视频生成从单次出片走向多模态创作引擎
发布时间:
Seedance 2.0:视频生成从单次出片走向多模态创作引擎
MeanFlow:一步生成不是蒸馏,而是学习平均速度场
发布时间:
MeanFlow:一步生成不是蒸馏,而是学习平均速度场
Recommended citation: Zhengyang Geng et al. Mean Flows for One-step Generative Modeling. NeurIPS 2025.
Download Paper
Mean Mode Screaming:为什么 1000 层 Diffusion Transformer 会被 token 均值拖垮
发布时间:
Mean Mode Screaming:为什么 1000 层 Diffusion Transformer 会被 token 均值拖垮
Recommended citation: Pengqi Lu. Mean Mode Screaming: Mean--Variance Split Residuals for 1000-Layer Diffusion Transformers. arXiv:2605.06169, 2026.
Download Paper
Flow-OPD:把多任务奖励对齐改写成 Flow Matching 的 on-policy 蒸馏
发布时间:
Flow-OPD:把多任务奖励对齐改写成 Flow Matching 的 on-policy 蒸馏
Recommended citation: Zhen Fang et al. Flow-OPD: On-Policy Distillation for Flow Matching Models. arXiv:2605.08063, 2026.
Download Paper
Edit2Restore:把图像复原改写成少样本图像编辑
发布时间:
Edit2Restore:把图像复原改写成少样本图像编辑
Recommended citation: Yılmaz et al., Edit2Restore: Few-Shot Image Restoration via Parameter-Efficient Adaptation of Pre-trained Editing Models, arXiv 2026
Download Paper
Code as Agent Harness:把代码看成 Agent 的运行底座
发布时间:
Code as Agent Harness:把代码看成 Agent 的运行底座
Recommended citation: Ning et al., Code as Agent Harness: Toward Executable, Verifiable, and Stateful Agent Systems, arXiv:2605.18747, 2026.
Download Paper
AsymFlow:把 latent flow 拉回 pixel space 的低秩速度参数化
发布时间:
AsymFlow:把 latent flow 拉回 pixel space 的低秩速度参数化
SHARP:单张照片在一秒内变成可实时渲染的 3D Gaussian 场
发布时间:
SHARP:单张照片在一秒内变成可实时渲染的 3D Gaussian 场
Recommended citation: Mescheder et al., Sharp Monocular View Synthesis in Less Than a Second, ICLR 2026
Download Paper
Follow the Mean:把参考样本变成 Flow Matching 的控制信号
发布时间:
Follow the Mean:把参考样本变成 Flow Matching 的控制信号
Recommended citation: Pedro M. P. Curvo, Maksim Zhdanov, Floor Eijkelboom, and Jan-Willem van de Meent. Follow the Mean: Reference-Guided Flow Matching. arXiv:2605.10302, 2026.
Download Paper
LPM 1.0:从 talking head 到实时对话角色 Performance Model
发布时间:
LPM 1.0:从 talking head 到实时对话角色 Performance Model
Recommended citation: Ailing Zeng et al. LPM 1.0: Video-based Character Performance Model. arXiv:2604.07823v2, 2026.
Download Paper
Self-Flow:把表征学习塞回 Flow Matching 训练目标里
发布时间:
Self-Flow:把表征学习塞回 Flow Matching 训练目标里
Recommended citation: Hila Chefer, Patrick Esser, Dominik Lorenz, Dustin Podell, Vikash Raja, Vinh Tong, Antonio Torralba, Robin Rombach. Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis. arXiv:2603.06507, 2026.
Download Paper
为什么 RoPE 对外推友好
发布时间:
RoPE 外推友好性完整解析
Self-Forcing 到 Self-Forcing++:让自回归视频扩散按推理方式训练
发布时间:
Self-Forcing 到 Self-Forcing++:让自回归视频扩散按推理方式训练
noao-vlm-2 数据集与评估系统分析
发布时间:
数据集与评估系统分析
noao-vlm-1 架构详细分析
发布时间:
nanoVLM 模型架构与数据流转分析
noao-vlm-0 train.py 详细分析
发布时间:
train.py 详细分析
noao-chat-7-nonochat 多卡训练指南
发布时间:
LLM 多卡训练完全指南
noao-chat-6-训练评估指南
发布时间:
LLM 模型评估验证完全指南
noao-chat-5-训练四阶段数据报告
发布时间:
nanochat 项目四阶段训练数据完全报告
noao-chat-4-rl阶段训练
发布时间:
LLM RL 训练完整解析
noao-chat-3-sft阶段训练
发布时间:
LLM SFT 训练完整解析
noao-chat-2-mid阶段训练
发布时间:
nonochat - LLM Mid 训练完整解析
noao-chat-1-base阶段训练
发布时间:
nonochat - LLM Base 训练完整解析
noao-chat-0-项目总体介绍
发布时间:
nanochat 项目深度分析
Efficient Rectified Flow for Image Fusion
发布时间:
[Paper Reading] Efficient Rectified Flow for Image Fusion(RFfusion)论文解读
3Blue1Brown 线性代数笔记
发布时间:
3Blue1Brown 线性代数笔记(几何直觉)
线性代数的常见概念集合
发布时间:
一、向量与向量空间相关(定义 + 几何直觉)
DiT4SR: Taming Diffusion Transformer for Real-World Image Super-Resolution
发布时间:
[Paper Reading] 基于 Diffusion Transformer 的真实世界超分辨率方法 DiT4SR
Dual Prompting Image Restoration with Diffusion Transformers
发布时间:
[Paper Reading] 基于扩散 Transformer 的双重提示图像复原 (DPIR)
FLOAT 深读:为什么 talking portrait 应该先生成 motion latent
发布时间:
FLOAT 论文精读:把 audio-driven talking portrait 的生成目标从 pixel video 换到 motion latent trajectory,再用 Flow Matching 快速采样。
Recommended citation: Ki, Min, and Chae. FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking Portrait. ICCV 2025.
Download Paper
Stable Video-Driven Portraits
发布时间:
[Paper Reading]:Stable Video-Driven Portraits — 基于 DiT 的高保真视频驱动人像生成
moco 论文摘要
发布时间:
MoCo: Momentum Contrast for Unsupervised Visual Representation Learning
小于1000的正整数立方和pair
发布时间:
找出所有满足 \(a^3+b^3=c^3+d^3\)的小于1000的正整数组合
什么是deep learning
发布时间:
前言
portfolio
Portfolio item number 1
Short description of portfolio item number 1
Portfolio item number 2
Short description of portfolio item number 2 
publications
Deep Cube-Pair Network for Hyperspectral Imagery Classification
Published in remotesensing, 2018
Improving Hyperspectral Image Classification with Unsupervised Knowledge Learning
Published in (IGARSS) IEEE International Geoscience and Remote Sensing Symposium, 2019
Learning Discriminative Compact Representation for Hyperspectral Imagery Classification
Published in IEEE Transactions on Geoscience and Remote Sensing (TGRS) , 2019
The Second Challenge on Real-World Face Restoration at NTIRE 2026: Methods and Results
Published in arXiv preprint arXiv:2604.10532v2 / CVPR 2026 Workshop, 2026
IConFace: Fine-Grained Identity Conditioning for Reference-Aware Face Restoration
Published in arXiv preprint arXiv:2605.02814, 2026
talks
Talk 1 on Relevant Topic in Your Field
发布时间:
This is a description of your talk, which is a markdown file that can be all markdown-ified like any other post. Yay markdown!
teaching
Teaching experience 1
Undergraduate course, University 1, Department, 2014
This is a description of a teaching experience. You can use markdown like any other post.
