# Is there any way to get round the overhead of numba prange with one thread?

**URL:** <https://numba.discourse.group/t/is-there-any-way-to-get-round-the-overhead-of-numba-prange-with-one-thread/3085>\
**Category:** Community Support\
**Created:** [November 29, 2025, 10:23pm UTC](https://numba.discourse.group/t/is-there-any-way-to-get-round-the-overhead-of-numba-prange-with-one-thread/3085 "2025-11-29T22:23:32Z")\
**Posts on this page:** 7\
**Page:** 1

<div class="post-metadata">

**Author:** ![lesshaste](https://yyz2.discourse-cdn.com/free1/user_avatar/numba.discourse.group/lesshaste/32/808_2.png) [@lesshaste](https://numba.discourse.group/u/lesshaste)\
**Post date:** [November 29, 2025, 10:23pm UTC](https://numba.discourse.group/t/is-there-any-way-to-get-round-the-overhead-of-numba-prange-with-one-thread/3085/1 "2025-11-29T22:23:32Z")

</div>

I like to be about to set the number of threads for my code to 1 sometimes. Unfortunately, all the numba code that uses prange then runs more slowly than if I had used range. Is there a way round this?

Would a GitHub issue about this be welcome? Maybe it would be possible for numba to skip the parallel machinery entirely if num threads equals 1?

---

<div class="post-metadata">

**Author:** ![nelson2005](https://yyz2.discourse-cdn.com/free1/user_avatar/numba.discourse.group/nelson2005/32/47_2.png) [@nelson2005](https://numba.discourse.group/u/nelson2005)\
**Post date:** [November 30, 2025, 5:21am UTC](https://numba.discourse.group/t/is-there-any-way-to-get-round-the-overhead-of-numba-prange-with-one-thread/3085/2 "2025-11-30T05:21:26Z")

</div>

> [@lesshaste](#):
>
> Unfortunately, all the numba code that uses prange then runs more slowly than if I had used range.

Do you have a minimal reproducer for the slowdown you’re seeing?

---

<div class="post-metadata">

**Author:** ![lesshaste](https://yyz2.discourse-cdn.com/free1/user_avatar/numba.discourse.group/lesshaste/32/808_2.png) [@lesshaste](https://numba.discourse.group/u/lesshaste)\
**Post date:** [December 1, 2025, 8:26am UTC](https://numba.discourse.group/t/is-there-any-way-to-get-round-the-overhead-of-numba-prange-with-one-thread/3085/3 "2025-12-01T08:26:38Z")

</div>

Here is an example:

```auto
import numpy as np
import numba as nb
from numba import njit, prange

nb.set_num_threads(1)

@njit(parallel=True, fastmath=True)
def lower_tri_matvec_par(L, x):
    """
    Multiply a lower triangular matrix L (n x n)
    by a vector x (length n).

    Returns y = L @ x, using only the lower triangle.
    """

    n = L.shape[0]
    y = np.zeros(n, dtype=L.dtype)

    for i in prange(n):
        s = 0.0
        # only sum up to the diagonal (j <= i)
        for j in range(i + 1):
            s += L[i, j] * x[j]
        y[i] = s
    return y

@njit(fastmath=True)
def lower_tri_matvec(L, x):
    """
    Multiply a lower triangular matrix L (n x n)
    by a vector x (length n).

    Returns y = L @ x, using only the lower triangle.
    """
    n = L.shape[0]
    y = np.zeros(n, dtype=L.dtype)
    for i in range(n):
        s = 0.0
        # only sum up to the diagonal (j <= i)
        for j in range(i + 1):
            s += L[i, j] * x[j]
        y[i] = s
    return y

N = 2000

L = np.tril(np.random.rand(N, N).astype(np.float32))
x = np.random.rand(N).astype(np.float32)

In [103]: %timeit -n 1000 -r 100 lower_tri_matvec_par(L, x)
851 μs ± 35 μs per loop (mean ± std. dev. of 100 runs, 1,000 loops each)

In [104]: %timeit -n 1000 -r 100 lower_tri_matvec(L, x)
819 μs ± 20.7 μs per loop (mean ± std. dev. of 100 runs, 1,000 loops each)

```

---

<div class="post-metadata">

**Author:** ![nelson2005](https://yyz2.discourse-cdn.com/free1/user_avatar/numba.discourse.group/nelson2005/32/47_2.png) [@nelson2005](https://numba.discourse.group/u/nelson2005)\
**Post date:** [December 1, 2025, 9:35pm UTC](https://numba.discourse.group/t/is-there-any-way-to-get-round-the-overhead-of-numba-prange-with-one-thread/3085/4 "2025-12-01T21:35:26Z")

</div>

Thanks… is it possible that compiling time is being included in your observations?

What is the result of you run %timeit a second time immediately after what you have now?

---

<div class="post-metadata">

**Author:** ![lesshaste](https://yyz2.discourse-cdn.com/free1/user_avatar/numba.discourse.group/lesshaste/32/808_2.png) [@lesshaste](https://numba.discourse.group/u/lesshaste)\
**Post date:** [December 2, 2025, 9:07am UTC](https://numba.discourse.group/t/is-there-any-way-to-get-round-the-overhead-of-numba-prange-with-one-thread/3085/5 "2025-12-02T09:07:45Z")

</div>

> [@lesshaste](#):
>
> `nb.set_num_threads(1)`

In [2]: %timeit -n 1000 -r 100 lower\_tri\_matvec\_par(L, x)  
831 μs ± 218 μs per loop (mean ± std. dev. of 100 runs, 1,000 loops each)

In [3]: %timeit -n 1000 -r 100 lower\_tri\_matvec(L, x)  
790 μs ± 18.3 μs per loop (mean ± std. dev. of 100 runs, 1,000 loops each)

In [4]: %timeit -n 1000 -r 100 lower\_tri\_matvec\_par(L, x)  
812 μs ± 3.56 μs per loop (mean ± std. dev. of 100 runs, 1,000 loops each)

In [5]: %timeit -n 1000 -r 100 lower\_tri\_matvec(L, x)  
803 μs ± 16.7 μs per loop (mean ± std. dev. of 100 runs, 1,000 loops each)

---

<div class="post-metadata">

**Author:** ![nelson2005](https://yyz2.discourse-cdn.com/free1/user_avatar/numba.discourse.group/nelson2005/32/47_2.png) [@nelson2005](https://numba.discourse.group/u/nelson2005)\
**Post date:** [December 3, 2025, 3:56am UTC](https://numba.discourse.group/t/is-there-any-way-to-get-round-the-overhead-of-numba-prange-with-one-thread/3085/6 "2025-12-03T03:56:39Z")

</div>

Thanks, your results are directionally similar to mine, though you clearly have a much faster machine than I do 😀

I don’t know of any way to avoid the ~1% slowdown you’re seeing due to the parallelization overhead. As I understand it numba doesn’t make any different code path for different thread counts, the compilation is binary. (Threaded or not threaded) So clearly I haven’t added much value with respect to your original post. ☹

---

<div class="post-metadata">

**Author:** ![Oyibo](https://avatars.discourse-cdn.com/v4/letter/o/4af34b/32.png) [@Oyibo](https://numba.discourse.group/u/Oyibo)\
**Post date:** [December 3, 2025, 9:11pm UTC](https://numba.discourse.group/t/is-there-any-way-to-get-round-the-overhead-of-numba-prange-with-one-thread/3085/7 "2025-12-03T21:11:16Z")

</div>

Hey @lesshaste ,

If disc space and compilation time is not critical, you can compile the same core algorithm twice, once with parallel=False and once with parallel=True.  
You can use a dispatcher to choose between them at runtime. This lets you switch modes within the same Python session without changing thread settings.

Here is an example:

```auto
import numpy as np
from numba import njit, prange

NBCONFIG = {'fastmath': True, 'cache': False}

@njit(**NBCONFIG)
def lower_tri_matvec(L, x):
    """Benchmark"""
    n = L.shape[0]
    y = np.zeros(n, dtype=L.dtype)
    for i in range(n):
        s = 0.0
        for j in range(i + 1):
            s += L[i, j] * x[j]
        y[i] = s
    return y

@njit(**NBCONFIG, inline='always')
def lower_tri_matvec_core(L, x):
    """Core algorithm to be inlined."""
    n = L.shape[0]
    # initialization overhead: use empty array instead
    y = np.zeros(n, dtype=L.dtype)  
    for i in prange(n):
        # casting overhead: use type aware scalar variable instead
        s = 0.0  
        for j in range(i + 1):
            s += L[i, j] * x[j]
        y[i] = s
    return y

@njit(**NBCONFIG, parallel=False)
def lower_tri_matvec_seq(L, x):
    """Sequential algorithm (same as lower_tri_matvec)."""
    return lower_tri_matvec_core(L, x)

@njit(**NBCONFIG, parallel=True)
def lower_tri_matvec_par(L, x):
    """Parallel algorithm."""
    return lower_tri_matvec_core(L, x)

@njit(**NBCONFIG)
def lower_tri_matvec_dispatch(L, x, parallel=False):
    return lower_tri_matvec_par(L, x) if parallel else lower_tri_matvec_seq(L, x)

N = 2000
L = np.tril(np.random.rand(N, N).astype(np.float32))
x = np.random.rand(N).astype(np.float32)

# warm-up
lower_tri_matvec(L, x)
lower_tri_matvec_dispatch(L, x)
lower_tri_matvec_dispatch(L, x, parallel=True)

%timeit -n 300 -r 100 lower_tri_matvec(L, x)
%timeit -n 300 -r 100 lower_tri_matvec_dispatch(L, x, parallel=False)
%timeit -n 300 -r 100 lower_tri_matvec_dispatch(L, x, parallel=True)

# 811 μs ± 80 μs per loop (mean ± std. dev. of 100 runs, 300 loops each)
# 793 μs ± 35.9 μs per loop (mean ± std. dev. of 100 runs, 300 loops each)
# 670 μs ± 29.8 μs per loop (mean ± std. dev. of 100 runs, 300 loops each)

# 794 μs ± 40 μs per loop (mean ± std. dev. of 100 runs, 300 loops each)
# 788 μs ± 22.3 μs per loop (mean ± std. dev. of 100 runs, 300 loops each)
# 673 μs ± 39.4 μs per loop (mean ± std. dev. of 100 runs, 300 loops each)

# 788 μs ± 24.6 μs per loop (mean ± std. dev. of 100 runs, 300 loops each)
# 794 μs ± 38 μs per loop (mean ± std. dev. of 100 runs, 300 loops each)
# 705 μs ± 112 μs per loop (mean ± std. dev. of 100 runs, 300 loops each)

```
