The Xuantie C950 processor demonstrates a decoupled Tensor Processing Unit (TPE) design that enhances AI computing capabilities without breaking compatibility with general-purpose computing, achieving 80 tokens per second on Qwen 3 (30B parameters) and 18 tokens per second on DeepSeek (671B parameters) while maintaining 85-90% utilization across matrix sizes and data types.
Deep Dive
Prerequisite Knowledge
- No data available.
Where to go next
- No data available.
Deep Dive
Xuantie | Demo Theatre | RISC-V Summit Europe 2026
Added:Good afternoon, everyone. My name is Guoren from Alibaba Damo Academy, Xuantie processor team. Today, I will talk about the 10 minutes to explore how RISC-V is accelerating the future computing and let me introduce our latest achievement, newest product, C950.
Xuantie RISC-V is a key achievement of Alibaba Damo Damo Academy. We are here to achieve as foundation and have built full scenario computing capabilities spanning from the device to the servers.
Uh Xuantie Flex series and the Xuantie RISC-V processor product lines such as the E series, C series, R series, and the Xuantie AI platform. These three pillars jointly support the Xuantie full scenario computing capabilities from edge AI devices to high-performance servers.
Xuantie has been widely adopted in the AI glasses, cameras, toys, robots, accelerator cards, boxes, and the various cloud servers. We currently have 900 authorized license and a 4.5 million billion chips shipped, about 400 authorized customers. Together, we are building a new era computing for the AI for everything.
Xuantie C950 50 one of the highest performance RISC-V CPU IP.
Its core microarchitecture highlights include the eight-wide instruction decoder and a 16-stage pipeline.
Out-of-order run windows exceeding 1,000 instructions.
Target frequency uh 3.2 GHz with a single core 70 scores of SPEC int 2006.
This is a truly server-grade high-performance RISC-V core.
C950 fully supports the RVA23 profile and includes a rich extensions.
AVW MO and the TSO memory model dynamically could a switch between each other. Vector vector crypto and vector length is about a 256 bits width and the with actual extension uh advanced interrupt architecture IOMMU and other extensions.
And the deep out of order execution plus new generation perfect engine the tensor processing unit up to eight tops which natively supported different and multi uh data formats such as FP4, FP8, FP FP16, BF16, INT8.
And the C950 has a three level heat cache hierarchy with the uh private L2 and the distributed high bandwidth L3.
And uh it supports mainstream interconnected for the G.E.
version and the G.F. uh uh f issue.
And the with the XI4 it's very easy to integrate with your knock, with your crossbar, your SOC design.
Many customers want to enhance the AI processing capability of their general purpose computing chips.
While matching their existing memory bandwidth typically several hundred gigabytes bytes per second such as DDR5 and the LPDDR4 could provide.
So, how do we meet this requirement?
One common approach is simply enlarge the vector unit and increase vector length.
However, this creates serious problems.
Unlike X86 and ARM which have additional SIMD instruction sets such as uh uh Neo such as the SSE.
RISC-V relies relies heavily on the vector.
If we make the vector unit very large, for example, over the billions over maybe 1,000 bit weights, the area and the and the cost of the per core would increase significantly. More importantly, vector context switch would become extremely expensive because the vector is now widely used in our kernel in both user space and the kernel space.
The switch happened everywhere.
For general computing chips, keeping the vector at a reasonable size is a critical for the system efficiency at the Linux kernel scheduling.
If we design the vector units different between the big cores and the little cores, they cannot be scheduled together under the Linux SMP.
This would have forced us to disable vector entirely entirely, which is highly highly undesirable for the developers.
So, what is our solution?
Keep the vector unit unified and the use a decoupled design to enhance the computing power without breaking the compatibility.
C950 uses an attached tensor processing unit or we call it TPE.
This is decoupled design does not consume consume the resources inside of the main core.
Uh and the one TPE can be shared across the multi C950 cores while maintaining the architectural consistency. Our TPE includes a powerful matrix unit and a vector unit. It can also connect to customers' own core processor and a customer's MPU.
In addition, TPE has a large uh 4,000 uh bit tail tail tail registers cache to maximize the data reuse and efficiency.
Now, let's look at the actual result on the right side.
Uh with the 16 uh TPEs, the C950 delivers 80 tokens per second on the Qwen 3, and the it's a 30 billion parameter with the active 3 billion active parameters.
Uh and the it's called uh give you a very low latency for the time to first the token.
On the much larger such as the Qwen 3, uh it could provide us a 64 tokens per second and a 1.7 seconds of TTFT.
Event even for the large very large model such as the DeepSeek uh uh 671 billion parameter, we can still reach uh 18 tokens per second.
Um using affordable DDR5 and the LPDDR4 memory. This memory very cheap, and uh we can use this cheap memory to run the large memory model.
These results demonstrate that the C950, which is attached the TP design, can efficiently run large language models on on the general-purpose computing platforms while fully leveraging the existing memory bandwidth.
Besides TPE, we also offer 10 uh Titan, which features a 4K ultra-wide vector engine for dedicated hardware accelerators. Both the solution support the unifying the addressing.
C950's tensor processing unit maintains, you can see the 85 from 85 to 89 90% utilization across different maximum matrix sizes and the data types.
In in in typical operators it delivers up to almost the four times speed up compared to the Revo products.
C910 C950's tensor processor unit Okay.
We okay.
This this is about of the our software ecosystem. We are building a full stack software ecosystem from the foundation to application. We started with the adoption and the porting in the early phase and then and now moving into the deep integration created real competitive competitive advantages in both application and AI.
Many customers already have their own MPU at all matrix computer units. In the AI era they want to maintain the computing Soviet Soviet sovereignty and work closely with our high performance processors.
To support this we we introduce the XuanTie flex series. It empowers customers with a full rank developer capability for RISC-V core based accelerators. We could we could provide the model for your verification and we have the develop based on the Gemmini it is a very common and we have developer in environment and we could provide a core DSA demo for the verification and we have the tool chain debugger compiler assembler and you you just write a some adjacent and some cell instruction description and all the tool chain will be generated. So, the customer just focus their application and connect to our customers uh MPU and their accelerator.
We are launch We are launching our next generation high-performance RISC-V processor series for data centers. And the lineup of processors from the C920 25 for the ultra-high efficiency to the C930 for the balanced performance and culminates in the C950, our ultra-high performance flagship chip processor.
These three flagship processors are built on our largest innovative microarchitecture.
Uh these uh are our massive product produced products from our customers that are already in the market.
They cover AI influence accelerator, intelligent security control chip, and SSD controllers, and the network switch.
Thank Thank you for your attention for in this AI era. Let's accelerate the future of the computing together with the RISC-V. Thank you very much.
Related Videos

Setting up a curved screen with Immersive Calibration Pro 4 and multiple cameras (P3D v4)
FlyerOneZero
23K views•2019-07-21

Robot Learning with Sparsity and Scarcity
allenai
379 views•2025-10-14

Jorge Mendez-Mendez: Unlocking Lifelong Robot Learning With Modularity (2023-10-05)
umassmlfl
237 views•2024-01-06

Northwestern’s MS in Robotics: Student Robotics Projects, 2023
NorthwesternEngineering
1K views•2024-05-31

"Perfect" Turns: Turning by the Gyro - FIRST LEGO League (FLL) SPIKE Prime + EV3 RePlay Programming
ZacharyTrautwein
94K views•2020-10-02

Gorkem Secer: TSLIP-based Deadbeat Running Control of Bipedal Robot ATRIAS
DynamicWalking-wv6qm
298 views•2018-06-22

Self-Driving Cars Need Lessons On Human Drivers | Maddie About Science
skunkbear
26K views•2018-08-21

Milrem Robotics’ THeMIS UGVs used in a live-fire manned-unmanned teaming exercise
MilremRobotics
99K views•2021-05-20
Trending

Playstation NO DISC/NO BUY Fight Is Over...
DavidJaffeGames
4K views•2026-07-23

Steam and Xbox Just Dropped The Hammer On PlayStation
OhNoItsAlexx
9K views•2026-07-23

Americans Confused in Australia for 17 Minutes Straight
IWrocker
17K views•2026-07-23

SuperBike Factory Has Gone... What's Next for the Motorcycle Industry?
thatbikersimon
11K views•2026-07-22