What Are CUDA Cores? How NVIDIA GPU Cores Work
Nerdivation · Published

A graphics card’s specifications might list thousands of CUDA cores. That sounds impressive, especially beside a CPU with eight or sixteen cores.
But what are those thousands of cores actually doing?
CUDA cores are arithmetic execution resources inside NVIDIA GPUs. In GeForce specifications, the advertised count refers to FP32-capable arithmetic lanes, not thousands of independent processors equivalent to CPU cores. NVIDIA’s architecture diagrams distinguish these resources from the larger structures that contain and coordinate them. [1]
Their count describes part of the hardware. It does not tell you how fast every application will run.
If you have already read CPU Cores vs GPU Cores: Why the Numbers Don’t Compare, this is the next step: understanding what NVIDIA is counting.
A CUDA Core Performs Arithmetic
GPUs repeatedly perform mathematical operations on large collections of numbers.
An image effect might adjust millions of color values. A simulation might calculate forces acting on particles. A numerical application might transform a large array of measurements.
CUDA cores provide arithmetic resources for this work.
FP32, also called single precision, is a 32-bit floating-point format. Floating-point formats represent numbers across a wide range of magnitudes, including fractional values.
However, “CUDA-core count” is not a count of every execution resource inside the GPU. Load/store units, special-function units, Tensor Cores, and RT Cores perform other kinds of work. A thread’s memory-load instruction, for example, uses memory-execution resources rather than turning a CUDA arithmetic lane into a complete processor. [2]
The GPU combines those resources to execute a program.
Where CUDA Cores Live
CUDA cores are organized inside larger structures called streaming multiprocessors, usually shortened to SMs.
An SM contains execution resources and supporting hardware for running GPU threads. Registers hold working values, scheduling hardware selects instructions to issue, and memory facilities help supply the data those instructions need. The exact arrangement varies by architecture. [3]
Think of an SM as a workshop.
The CUDA cores are some of its calculation tools. Registers keep current materials close at hand. Scheduling hardware decides which ready job can use the available equipment.
Counting the tools does not describe the entire workshop. You also need to know how work reaches them and whether they spend their time calculating or waiting.
CUDA Cores Are Not CUDA Threads
A thread is a logical instance of a program’s work. A CUDA core is a physical execution resource.
A program can launch far more threads than the GPU has CUDA cores. Those threads are scheduled onto the available hardware; they are not permanently paired with individual arithmetic lanes.
In NVIDIA’s programming model, threads belong to thread blocks. Within those blocks, threads are organized into warps of 32.
A kernel is a function launched to run on the GPU. Its threads execute the kernel code while operating on their own data. [3]
An image-processing kernel might assign different pixels to different threads. An array-processing kernel might assign different elements.
The number of threads describes how software divides its work. The CUDA-core count describes part of the hardware available to execute it.
How the GPU Keeps Its Arithmetic Units Busy
More execution resources help only when useful work reaches them.
Suppose one warp needs data that has not arrived yet. An SM’s scheduler can issue instructions from another ready warp while the first waits.
This helps hide memory latency. It does not eliminate the delay; it overlaps that delay with useful work.
Registers, shared memory, dependencies, and the available parallelism limit how much ready work the GPU can maintain. When no suitable warp is ready, arithmetic resources can sit idle. [4]
An enormous workshop is not productive when every tool is waiting for materials.
What Happens When Threads Take Different Paths?
Imagine an image effect that treats bright pixels differently from dark pixels.
Threads in the same warp may follow different branches. This is called branch divergence.
When executing a path, threads that do not participate are inactive. Processing the different paths can therefore use fewer arithmetic lanes than processing a path shared by the whole warp. [4]
NVIDIA introduced independent thread scheduling with Volta. It gives threads within a warp more flexible execution and synchronization behavior, but it does not turn CUDA cores into independent CPU-like processors or remove the efficiency consequences of divergent work. [5]
CUDA Cores, Tensor Cores, and RT Cores
These names describe different hardware responsibilities.
CUDA cores provide arithmetic resources used throughout GPU computation.
Tensor Cores specialize in supported matrix multiply-and-accumulate operations. Those operations are important in neural networks and other numerical workloads. Their supported formats and throughput vary by generation, and software must use a compatible execution path to benefit from them. [6]
RT Cores, found in NVIDIA RTX GPUs, accelerate particular ray-tracing operations, including traversal of scene acceleration structures and ray–triangle intersection tests. They do not perform every step needed to produce a ray-traced image. [2]
Term | What it describes |
|---|---|
CUDA core | An arithmetic execution resource; GeForce counts describe FP32-capable lanes |
Streaming multiprocessor | A larger processing structure containing execution and supporting resources |
CUDA thread | A logical instance of a kernel’s work |
Warp | A group of 32 NVIDIA GPU threads |
Tensor Core | Hardware specializing in supported matrix operations |
RT Core | Hardware accelerating particular ray-tracing operations on RTX GPUs |
A CUDA application can use Tensor Cores. “CUDA” also names NVIDIA’s computing platform and programming model; it does not mean every instruction in a CUDA program runs on CUDA cores. [3, 6]
A Concrete Example: The RTX 4090
The GeForce RTX 4090 has 128 SMs, each containing 128 CUDA cores, for a total of:
128 × 128 = 16,384 CUDA cores
That does not mean it contains 16,384 independent CPU-like processors. Its SMs coordinate threads and issue work to their execution resources. NVIDIA’s Ada documentation separately lists the card’s Tensor and RT resources. [2]
The large number describes arithmetic capacity distributed across a coordinated architecture.
Why Comparisons Across Generations Are Difficult
The hardware surrounding a CUDA core changes between generations.
Volta added dedicated INT32 execution resources alongside its FP32 resources, allowing integer and floating-point operations to execute concurrently. Turing carried that approach into GeForce GPUs. [5, 7]
The GA10x Ampere design provides a particularly useful comparison.
A Turing SM has 64 FP32 CUDA cores alongside separate integer resources. GA10x expands the FP32-capable resources to 128 per SM. Part of its execution hardware can process either FP32 or INT32 work, so the instruction mix affects how the available capacity is used.
This was a real increase in FP32 capability, not merely a new label for unchanged hardware. But doubling that capability does not guarantee twice the performance in a workload mixing arithmetic, memory operations, and other instructions. [1]
The qualification matters: GA10x is not every Ampere GPU. NVIDIA’s GA100 data-center design has a different arrangement.
Does Twice the CUDA-Core Count Mean Twice the Speed?
No.
Two limitations help explain why.
Memory-bound work spends relatively more time moving data than performing calculations. Additional arithmetic resources may provide little benefit when data delivery is the bottleneck.
Compute-bound work performs enough calculation for execution throughput to become the limiting factor. More capable arithmetic hardware can help when the application supplies enough suitable parallel work. [8]
An application may also spend time on the CPU, transferring data, launching kernels, or synchronizing operations. Accelerating one part does not automatically accelerate the entire application by the same amount. [9]
CUDA-core count describes an ingredient in performance. It is not a benchmark result.
How Should You Use CUDA-Core Counts?
Treat the number as supporting information.
Start with measurements that reflect what you actually want to do:
- Gaming: Compare performance in the games, resolutions, and settings you use.
- Rendering: Compare results from your rendering engine, and check whether large scenes fit in VRAM.
- AI: Check model size, VRAM capacity, numerical precision, and software support.
- Scientific computing: Check the application and the arithmetic precision it requires.
Then use the specifications to understand the results.
A GPU with more CUDA cores might have greater arithmetic capacity but still be constrained by memory bandwidth, insufficient VRAM, or software that cannot keep its execution resources busy.
Understanding those limitations is more useful than choosing the largest number on a specification sheet.
The Practical Explanation
CUDA cores are arithmetic resources inside a larger, coordinated NVIDIA GPU architecture.
They matter because GPUs can apply many calculations across large datasets. Their usefulness depends on the streaming multiprocessors, memory system, software, and workload around them.
CUDA-core count tells you something about execution hardware. It does not tell you how many miniature CPUs a GPU contains or how fast every application will run.
For the broader architectural explanation, continue with Why GPUs Are Faster Than CPUs and When They Aren’t.
Sources
Source | Supports |
|---|---|
“GA10x SM Architecture”: FP32/INT32 execution resources and the comparison with Turing. | |
SM organization, distinct hardware resources, RT Core responsibilities, and RTX 4090 specifications. | |
GPU hardware organization, kernels, thread blocks, and “Warps and SIMT.” | |
“SIMT Execution Model” and “Hardware Multithreading”: divergence, warp scheduling, and latency hiding. | |
“Independent Thread Scheduling” and “Integer Arithmetic”: features introduced with Volta. | |
Tensor Core use and matrix-operation performance considerations. | |
“Streaming Multiprocessor” and “Integer Arithmetic”: Turing’s execution resources and relationship to Volta. | |
Compute-bound and memory-bound performance. | |
Whole-application performance, parallelization, and host/device data transfers. |