Tencent Is Reshaping Hunyuan’s Multimodal AI Team, Source Says(Yicai) Aug. 19 -- Tencent Holdings is making personnel changes to the multimodal team of its Hunyuan artificial intelligence arm, according to a source close to the Chinese internet giant, less than a month after merging the unit’s large language model and multimodal model departments.
The roles of Bo Liefeng, the former head of Hunyuan's multimodal department, will likely change, while a senior researcher who recently left Kuaishou Technology's Kling AI has returned to Tencent, but joined its gaming business, according to the source. Bo, who joined the Hunyuan team in July last year, primarily oversees the multimodal business.
On July 24, Tencent announced that it was combining Hunyuan’s LLM and multimodal model departments, creating a single foundation model department led by the Shenzhen-based company’s Chief AI Scientist Yao Shunyu.
Tian Yonglong, a PhD graduate of the Massachusetts Institute of Technology who previously worked with Yao at OpenAI, joined the Hunyuan multimodal team early last month, while Zhong Zhao, head of Hunyuan's text-to-video and text-to-image algorithms, left the company. Responsible for developing vision-language models, Tian reports directly to Yao.
Tian may take on more responsibilities and become the de facto head of Hunyuan's multimodal business under Yao, the source pointed out.
Tencent has faced criticism for lagging domestic rivals such as Alibaba Group Holding, Baidu, and ByteDance in developing AI. Yao rejected that narrative last month, arguing that “AI is a long-term game” and that, in many respects, “the second half of the AI race is only just beginning.”
The company is making great strides in becoming an AI-empowered version of itself, founder Pony Ma said earlier this month after nearly tripling second-quarter capital expenditure from a year earlier so as crank up investment in AI.
The personnel changes follow Hunyuan’s limited progress in the multimodal market. Tencent launched the HunyuanImage-3.0 image generation model last September and the HunyuanVideo-1.5 video generation model last November, but has not released a new multimodal model since.
The HunyuanVideo-1.5 has 8.3 billion parameters but does not support end-to-end audio-visual synchronized generation, meaning that it cannot create a video with synchronized sound and visuals in a single step based solely on one prompt. It has also failed to solve the problem of synchronizing lip movements with speech.
Alibaba, TikTok-owner ByteDance, and Kling AI have each made breakthroughs in generating videos with synchronized audio and visuals and advanced their multimodal reasoning capabilities this year. For example, ByteDance has released successive updates to its Seedance 2.0 and Seedance 2.5 models.
Editor: Martin Kadiev
