Skip to content

ENH: use x86-simd-sort for descending sorts and partitions - #32690

Merged
ngoldbaum merged 5 commits into
numpy:mainfrom
eendebakpt:perf/x86-simd-descending-sort
Sep 21, 2026
Merged

ngoldbaum merged 5 commits into
numpy:mainfrom
eendebakpt:perf/x86-simd-descending-sort

Conversation

@eendebakpt

Copy link
Copy Markdown
Contributor

PR summary

The x86 SIMD paths were gated behind !reverse since gh-31345, because x86-simd-sort placed NaNs at the start of a descending result while NumPy keeps them at the end (see npy::cmp<Tag, reverse>). The vendored library gained a separate nans_last switch in numpy/x86-simd-sort#235, which came in with the submodule bump in gh-31947, so descending can now be forwarded directly: nans_last defaults to true and gives exactly NumPy's ordering.

Measured on AVX2 over 1M elements, descending now runs at the speed of ascending: ~4-9x faster than before for sort/partition and ~2-3x for argsort/argpartition.

AI Disclosure

Claude was used in creation of the PR. Identified while researching options for making stable sort the default.

The x86 SIMD paths were gated behind `!reverse` since numpygh-31345, because
x86-simd-sort placed NaNs at the start of a descending result while NumPy
keeps them at the end (see `npy::cmp<Tag, reverse>`).  The vendored library
gained a separate `nans_last` switch in numpy/x86-simd-sort#235, which came
in with the submodule bump in numpygh-31947, so `descending` can now be forwarded
directly: `nans_last` defaults to true and gives exactly NumPy's ordering.

Removes the dispatch guards in quicksort.hpp/selection.hpp and the
`assert(!reverse)` in the three dispatch translation units, and threads
`reverse` through QSelect/ArgQSelect so partition and argpartition are
covered as well.

Measured on AVX2 over 1M elements, descending now runs at the speed of
ascending: ~4-9x faster than before for `sort`/`partition` and ~2-3x for
`argsort`/`argpartition`.

Closes part of numpygh-31423.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
eendebakpt and others added 2 commits September 18, 2026 09:44
The `axis=0` example had a tied column (`[3, 3]`), so the indices it printed
depended on how argpartition happens to break ties.  With descending
argpartition now going through x86-simd-sort, which implements descending as
ascending followed by a reverse, that tie flips on x86 while the scalar path
used on other platforms keeps the old order -- the expected output would be
correct only on x86.

Drop the tie instead of pinning one platform's answer: a 2x4 array keeps the
"row and its reverse" structure of the example but has distinct values in
every column and row, so the result is uniquely determined by the values and
is the same for any implementation.  top_k already documents that the indices
it returns for duplicate values are not stable.

Co-Authored-By: Claude Opus 5 (1M context) <[email protected]>
@ikrommyd

ikrommyd commented Sep 18, 2026 •

Copy link
Copy Markdown
Member

Measured on AVX2 over 1M elements, descending now runs at the speed of ascending: ~4-9x faster than before for sort/partition and ~2-3x for argsort/argpartition.

Is there any chance you could share the benchmark? I'd be interested to run it on my ryzen cpu that supports avx512 as well.

Comment thread numpy/_core/fromnumeric.py
@eendebakpt

Copy link
Copy Markdown
Contributor Author

Measured on AVX2 over 1M elements, descending now runs at the speed of ascending: ~4-9x faster than before for sort/partition and ~2-3x for argsort/argpartition.

Is there any chance you could share the benchmark? I'd be interested to run it on my ryzen cpu that supports avx512 as well.

Results on my machine: i7-13650HX (AVX2, no AVX-512), Linux, CPython 3.14, release builds. The number from the OP is on a different machine.

function dtype asc (main) asc (PR) desc (main) desc (PR) desc speedup
sort float64 8.08 ms 8.01 ms 67.8 ms 8.06 ms 8.42x
argsort float64 30.3 ms 28.7 ms 88.4 ms 29 ms 3.05x
partition float64 1.32 ms 1.28 ms 9.47 ms 1.39 ms 6.80x
argpartition float64 2.41 ms 2.23 ms 10.7 ms 2.45 ms 4.38x
sort float32 4.62 ms 4.57 ms 68.5 ms 4.56 ms 15.00x
argsort float32 31.8 ms 28.6 ms 84.8 ms 28.8 ms 2.94x
partition float32 998 us 968 us 9 ms 934 us 9.63x
argpartition float32 2.18 ms 2.1 ms 10.7 ms 2.29 ms 4.66x
sort float16 64 ms 63.9 ms 65.2 ms 64.9 ms 1.00x
argsort float16 84.5 ms 84.3 ms 91.5 ms 91 ms 1.00x
partition float16 15.3 ms 15.2 ms 12.8 ms 12.8 ms 1.00x
argpartition float16 18.1 ms 18 ms 15.4 ms 15.3 ms 1.01x
sort int64 12.8 ms 12.8 ms 60.9 ms 12.7 ms 4.79x
argsort int64 37.6 ms 36.6 ms 81 ms 36.5 ms 2.22x
partition int64 1.36 ms 1.34 ms 7.31 ms 1.22 ms 6.02x
argpartition int64 2.8 ms 2.75 ms 8.68 ms 2.89 ms 3.00x
sort int32 4.79 ms 4.8 ms 62 ms 4.77 ms 13.00x
argsort int32 27.9 ms 27.8 ms 77.8 ms 28.1 ms 2.77x
partition int32 660 us 652 us 9.1 ms 651 us 13.98x
argpartition int32 1.74 ms 1.79 ms 10.7 ms 2.01 ms 5.33x
sort int16 54.8 ms 54.9 ms 55 ms 55.9 ms 0.98x
argsort int16 71.9 ms 72.5 ms 72.6 ms 72.3 ms 1.00x
partition int16 8.4 ms 8.36 ms 8.46 ms 8.47 ms 1.00x
argpartition int16 10.1 ms 10 ms 10.2 ms 10.2 ms 1.00x
n = 10,000
function dtype asc (main) asc (PR) desc (main) desc (PR) desc speedup
sort float64 48.1 us 47.8 us 428 us 48.4 us 8.85x
argsort float64 146 us 146 us 522 us 147 us 3.56x
partition float64 15.7 us 15.9 us 71.4 us 13.8 us 5.17x
argpartition float64 21.3 us 21.1 us 81.4 us 22.3 us 3.66x
sort float32 30.7 us 29.5 us 440 us 30.8 us 14.28x
argsort float32 227 us 161 us 639 us 162 us 3.95x
partition float32 10.5 us 10.9 us 70.7 us 10.3 us 6.88x
argpartition float32 23.2 us 22.8 us 82.5 us 24 us 3.44x
sort float16 528 us 526 us 518 us 518 us 1.00x
argsort float16 643 us 638 us 660 us 641 us 1.03x
partition float16 164 us 164 us 110 us 110 us 1.01x
argpartition float16 194 us 195 us 121 us 120 us 1.00x
sort int64 82.8 us 82.8 us 388 us 83.3 us 4.66x
argsort int64 206 us 207 us 476 us 208 us 2.28x
partition int64 12.9 us 13 us 89.2 us 19.5 us 4.56x
argpartition int64 25.5 us 25.8 us 105 us 27.5 us 3.83x
sort int32 29.3 us 29.4 us 393 us 28.9 us 13.59x
argsort int32 153 us 153 us 475 us 154 us 3.07x
partition int32 6.8 us 6.84 us 85.6 us 5.82 us 14.70x
argpartition int32 16.4 us 16.4 us 99.6 us 18 us 5.54x
sort int16 378 us 379 us 378 us 385 us 0.98x
argsort int16 464 us 464 us 466 us 459 us 1.01x
partition int16 45.1 us 45.5 us 81.9 us 80.1 us 1.02x
argpartition int16 53.5 us 52.5 us 94.2 us 93.2 us 1.01x
bench_descending_sort.py
"""
pyperf benchmark for numpy/numpy#32690: x86-simd-sort for descending
sort / argsort / partition / argpartition.

Run once per NumPy build and compare the results:

    python bench_descending_sort.py -o main.json      # in the env with main
    python bench_descending_sort.py -o pr.json        # in the env with the PR
    python -m pyperf compare_to main.json pr.json --table

Options (on top of the usual pyperf ones such as --fast, --rigorous,
--affinity):

    --dtypes float64,int32     dtypes to run (default: see DTYPES)
    --sizes 10000,1000000      array sizes to run (default: 1000000)
    --funcs sort,argsort       functions to run (default: all four)
    --kth-frac 0.5             kth for (arg)partition, as a fraction of n

All functions are called out-of-place (``np.sort(a)``, not ``a.sort()``) so
every call sees the same unsorted input.  That includes a copy of the input
(or the creation of an index array), which is the same on both builds; the
``copy`` benchmark measures that fixed cost so it can be subtracted.
"""
from functools import partial

import numpy as np
import pyperf

DTYPES = ["float64", "float32", "float16", "int64", "int32", "int16"]
SIZES = [1_000_000]
FUNCS = ["sort", "argsort", "partition", "argpartition"]


def make_data(dtype, n):
    rng = np.random.default_rng(12345)
    dtype = np.dtype(dtype)
    if dtype.kind == "f":
        return rng.standard_normal(n).astype(dtype)
    info = np.iinfo(dtype)
    return rng.integers(info.min, info.max, size=n, dtype=dtype, endpoint=True)


def add_cmdline_args(cmd, args):
    cmd.extend(["--dtypes", args.dtypes, "--sizes", args.sizes,
                "--funcs", args.funcs, "--kth-frac", str(args.kth_frac)])


def main():
    runner = pyperf.Runner(add_cmdline_args=add_cmdline_args)
    parser = runner.argparser
    parser.add_argument("--dtypes", default=",".join(DTYPES))
    parser.add_argument("--sizes", default=",".join(map(str, SIZES)))
    parser.add_argument("--funcs", default=",".join(FUNCS))
    parser.add_argument("--kth-frac", type=float, default=0.5)
    args = runner.parse_args()

    # Recorded in the JSON so it is clear which build and which SIMD
    # kernels produced the numbers (`pyperf metadata file.json`).
    from numpy._core._multiarray_umath import __cpu_features__
    runner.metadata["numpy_version"] = np.__version__
    runner.metadata["numpy_file"] = np.__file__
    runner.metadata["numpy_cpu_features"] = " ".join(
        name for name, have in __cpu_features__.items() if have)

    for dtype in args.dtypes.split(","):
        for n in map(int, args.sizes.split(",")):
            a = make_data(dtype, n)
            kth = min(n - 1, int(n * args.kth_frac))
            tag = f"{dtype} n={n}"

            runner.bench_func(f"copy {tag}", a.copy)
            for name in args.funcs.split(","):
                func = getattr(np, name)
                extra = (kth,) if "partition" in name else ()
                for order, descending in (("asc", False), ("desc", True)):
                    # bench_func() does not forward keyword arguments
                    runner.bench_func(
                        f"{name} {tag} {order}",
                        partial(func, a, *extra, descending=descending))


if __name__ == "__main__":
    main()

@ngoldbaum
ngoldbaum requested a review from r-devulap September 18, 2026 14:43

@ngoldbaum ngoldbaum left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I benchmarked this on an i5-8600K on Windows using an AI model. It supports AVX2 and X86_V3 but not AVX-512. I can confirm the performance improvement. I'd also like someone with an AVX-512 CPU to test. I have some nitpicks about the release note, see below.

use the same SIMD kernels as the ascending versions for most integer and
floating point dtypes. On AVX2 this is 2-9x faster than before. As a
result the indices returned for tied elements may differ from previous
versions; as before, they are not guaranteed.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The release note filename should use the number for this PR. I've noticed Claude is uncomfortable with the uncertainty of not knowing the PR number ahead of time so it randomly does other things.

The last sentence doesn't talk about stable sorts. Rather than making it more precise, you could also just not include the sentence about differing results for unstable sorts: we don't guarantee those.

@MaanasArora MaanasArora left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @eendebakpt, overall looks good to me! Just one real change inline. This is largely straightforward as the major update was in x86-simd-sort.

It would be nice to run the asv sort and partition benchmarks too, I think.

Comment thread numpy/_core/src/npysort/selection.hpp Outdated

template<typename Tag, typename T>
inline bool quickselect_dispatch(T* v, npy_intp num, npy_intp kth)
inline bool quickselect_dispatch(T* v, npy_intp num, npy_intp kth, bool reverse)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should this not be a template parameter as in quicksort_dispatch?

Comment thread numpy/_core/src/npysort/selection.hpp Outdated
---------------------------------------------------------------
`numpy.sort`, `numpy.argsort`, `numpy.partition` and `numpy.argpartition`
with ``descending=True`` no longer fall back to scalar code on x86, and now
use the same SIMD kernels as the ascending versions for most integer and

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Are there any exceptions (dtypes for which we don't use the same kernels)?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There are some dtypes where no kernels are used, but if used they are the same.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Right thanks, maybe something like this then? (But total nit)

... and now use the same SIMD kernels as the ascending versions for all dtypes that supported SIMD optimization.

or even

... and now use the same SIMD optimizations as with descending=False.

MaanasArora

This comment was marked as duplicate.

Make `reverse` a template parameter of quickselect_dispatch and
argquickselect_dispatch, consistent with the quicksort dispatchers.
Name the release note after the PR and drop the sentence on tie order.

Co-Authored-By: Claude Fable 5.1 <[email protected]>
@ikrommyd

Copy link
Copy Markdown
Member

I'm not getting that significant speed up on AMD Ryzen 7 7800X3D. I ran the benchmark you posted earlier by installing numpy like python -m pip install -vvv . -C'setup-args=-Dbuildtype=release' -C'build-dir=build-dir' --no-clean --no-build-isolation on main and on your PR branch. Then I ran python -m pyperf compare_to --table --table-format md main.json branch.json to compare the pyperf jsons. And this is the output:

Benchmark main branch
copy float64 n=1000000 138 us 137 us: 1.00x faster
sort float64 n=1000000 desc 56.8 ms 38.4 ms: 1.48x faster
argsort float64 n=1000000 asc 50.6 ms 50.2 ms: 1.01x faster
argsort float64 n=1000000 desc 75.1 ms 50.3 ms: 1.49x faster
partition float64 n=1000000 desc 7.21 ms 6.54 ms: 1.10x faster
argpartition float64 n=1000000 desc 9.10 ms 6.63 ms: 1.37x faster
sort float32 n=1000000 desc 56.4 ms 23.0 ms: 2.45x faster
argsort float32 n=1000000 desc 75.3 ms 47.9 ms: 1.57x faster
partition float32 n=1000000 desc 7.16 ms 4.45 ms: 1.61x faster
argpartition float32 n=1000000 asc 6.25 ms 6.26 ms: 1.00x slower
argpartition float32 n=1000000 desc 9.27 ms 6.35 ms: 1.46x faster
sort float16 n=1000000 asc 22.5 ms 22.4 ms: 1.00x faster
sort float16 n=1000000 desc 46.3 ms 22.9 ms: 2.02x faster
argsort float16 n=1000000 asc 58.5 ms 58.6 ms: 1.00x slower
argsort float16 n=1000000 desc 60.4 ms 60.3 ms: 1.00x faster
partition float16 n=1000000 desc 8.03 ms 3.92 ms: 2.05x faster
argpartition float16 n=1000000 asc 11.7 ms 11.7 ms: 1.00x faster
copy int64 n=1000000 137 us 139 us: 1.01x slower
sort int64 n=1000000 desc 39.8 ms 36.7 ms: 1.09x faster
argsort int64 n=1000000 desc 53.1 ms 48.4 ms: 1.10x faster
partition int64 n=1000000 desc 4.49 ms 6.30 ms: 1.40x slower
argpartition int64 n=1000000 asc 7.49 ms 7.49 ms: 1.00x slower
argpartition int64 n=1000000 desc 5.65 ms 7.57 ms: 1.34x slower
sort int32 n=1000000 desc 39.2 ms 22.2 ms: 1.77x faster
argsort int32 n=1000000 desc 52.1 ms 46.8 ms: 1.11x faster
partition int32 n=1000000 desc 5.60 ms 3.37 ms: 1.66x faster
argpartition int32 n=1000000 desc 7.25 ms 7.61 ms: 1.05x slower
copy int16 n=1000000 35.4 us 35.7 us: 1.01x slower
sort int16 n=1000000 desc 36.8 ms 15.9 ms: 2.31x faster
argsort int16 n=1000000 asc 47.8 ms 48.1 ms: 1.01x slower
argsort int16 n=1000000 desc 48.0 ms 48.3 ms: 1.01x slower
partition int16 n=1000000 desc 5.36 ms 3.03 ms: 1.77x faster
argpartition int16 n=1000000 asc 6.94 ms 6.93 ms: 1.00x faster
Geometric mean (ref) 1.14x faster

Benchmark hidden because not significant (21): sort float64 n=1000000 asc, partition float64 n=1000000 asc, argpartition float64 n=1000000 asc, copy float32 n=1000000, sort float32 n=1000000 asc, argsort float32 n=1000000 asc, partition float32 n=1000000 asc, copy float16 n=1000000, partition float16 n=1000000 asc, argpartition float16 n=1000000 desc, sort int64 n=1000000 asc, argsort int64 n=1000000 asc, partition int64 n=1000000 asc, copy int32 n=1000000, sort int32 n=1000000 asc, argsort int32 n=1000000 asc, partition int32 n=1000000 asc, argpartition int32 n=1000000 asc, sort int16 n=1000000 asc, partition int16 n=1000000 asc, argpartition int16 n=1000000 desc

@tylerjereddy

Copy link
Copy Markdown
Contributor

I was about to check on another AVX 512 arch on a local supercomputer (Intel(R) Xeon(R) Platinum 8480+). I'll move on though, since Ryzen should do it.

@ikrommyd

Copy link
Copy Markdown
Member

I was about to check on another AVX 512 arch on a local supercomputer (Intel(R) Xeon(R) Platinum 8480+). I'll move on though, since Ryzen should do it.

yours supports AVX512_SPR as well though which mine doesn't so if you wanna do it, go ahead.

@MaanasArora MaanasArora left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the change, LGTM! I don't think we need to block for more benchmarking, though it's nice to have; descending SIMD is something we need to support anyway.

(If the benchmarking results are unexpected, I guess we should compare first with the ascending versions, then check with x86-simd-sort if they do diverge.)

Edit: sorry one thing, let's ping @seberg, as we even planned to backport this as a performance bug which might still be nice?

@ikrommyd

ikrommyd commented Sep 19, 2026 •

Copy link
Copy Markdown
Member

Okay yeah I ran the benchmarks with NPY_DISABLE_CPU_FEATURES="X86_V4 AVX512_ICL AVX512_SPR" and apparently it's just a known thing that Zen 4 sorting with avx512 is just bad. I'm getting this below on avx2 with my cpu as well. I'm wondering if we can do something to automatically disable this for people on Zen 4 🤣. Dispatching to avx512 on Zen 4 feels wrong.

My AI model informs me that:

  • On Intel (Skylake-X, Ice Lake, Sapphire Rapids), x86-simd-sort's AVX-512 path is its fastest path, faster than AVX2. That's the hardware it was written and tuned for, at Intel.
  • On Zen 4, most AVX-512 instructions run fine, but vpcompressd/q/ps/pd with a memory destination is microcoded and very slow, reportedly over 100 cycles per instruction. The register-destination form of the same instruction is fast. x86-simd-sort's partition step does exactly this memory-destination compress-store on every vector (_mm512_mask_compressstoreu_*), so it hits the worst case.
  • Zen 5 reportedly fixed this.

Edit: apparently this is know and there are issues already open for it

So LGTM too! Not formally approving as I very quickly skimmed through the code but it looked like simple changes.

Benchmark results with AVX2 only
Benchmark main branch
copy float64 n=1000000 138 us 139 us: 1.01x slower
sort float64 n=1000000 asc 6.78 ms 6.84 ms: 1.01x slower
sort float64 n=1000000 desc 56.8 ms 6.80 ms: 8.36x faster
argsort float64 n=1000000 desc 75.1 ms 18.7 ms: 4.01x faster
partition float64 n=1000000 asc 1.01 ms 948 us: 1.07x faster
partition float64 n=1000000 desc 7.21 ms 1.06 ms: 6.82x faster
argpartition float64 n=1000000 asc 1.61 ms 1.63 ms: 1.01x slower
argpartition float64 n=1000000 desc 9.10 ms 1.75 ms: 5.21x faster
copy float32 n=1000000 67.7 us 68.2 us: 1.01x slower
sort float32 n=1000000 asc 4.08 ms 4.12 ms: 1.01x slower
sort float32 n=1000000 desc 56.4 ms 4.08 ms: 13.83x faster
argsort float32 n=1000000 asc 19.0 ms 19.3 ms: 1.02x slower
argsort float32 n=1000000 desc 75.3 ms 19.4 ms: 3.88x faster
partition float32 n=1000000 asc 697 us 765 us: 1.10x slower
partition float32 n=1000000 desc 7.16 ms 774 us: 9.25x faster
argpartition float32 n=1000000 desc 9.26 ms 1.75 ms: 5.30x faster
sort float16 n=1000000 asc 40.1 ms 40.4 ms: 1.01x slower
sort float16 n=1000000 desc 46.3 ms 45.7 ms: 1.01x faster
argsort float16 n=1000000 desc 60.4 ms 60.3 ms: 1.00x faster
argpartition float16 n=1000000 asc 11.8 ms 11.7 ms: 1.00x faster
sort int64 n=1000000 asc 7.78 ms 7.66 ms: 1.02x faster
sort int64 n=1000000 desc 39.8 ms 7.69 ms: 5.18x faster
argsort int64 n=1000000 desc 53.0 ms 20.1 ms: 2.64x faster
partition int64 n=1000000 asc 709 us 712 us: 1.00x slower
partition int64 n=1000000 desc 4.49 ms 651 us: 6.89x faster
argpartition int64 n=1000000 desc 5.65 ms 1.74 ms: 3.25x faster
copy int32 n=1000000 68.3 us 69.0 us: 1.01x slower
sort int32 n=1000000 asc 4.03 ms 4.00 ms: 1.01x faster
sort int32 n=1000000 desc 39.2 ms 3.73 ms: 10.51x faster
argsort int32 n=1000000 desc 52.1 ms 19.1 ms: 2.73x faster
partition int32 n=1000000 asc 437 us 438 us: 1.00x slower
partition int32 n=1000000 desc 5.60 ms 409 us: 13.69x faster
argpartition int32 n=1000000 asc 1.35 ms 1.34 ms: 1.00x faster
argpartition int32 n=1000000 desc 7.25 ms 1.45 ms: 4.99x faster
copy int16 n=1000000 35.4 us 35.7 us: 1.01x slower
sort int16 n=1000000 asc 36.7 ms 36.6 ms: 1.00x faster
sort int16 n=1000000 desc 36.8 ms 36.6 ms: 1.01x faster
argsort int16 n=1000000 asc 47.9 ms 48.3 ms: 1.01x slower
partition int16 n=1000000 asc 5.15 ms 5.21 ms: 1.01x slower
partition int16 n=1000000 desc 5.36 ms 5.28 ms: 1.01x faster
argpartition int16 n=1000000 asc 6.93 ms 6.92 ms: 1.00x faster
Geometric mean (ref) 1.69x faster

Benchmark hidden because not significant (13): argsort float64 n=1000000 asc, argpartition float32 n=1000000 asc, copy float16 n=1000000, argsort float16 n=1000000 asc, partition float16 n=1000000 asc, partition float16 n=1000000 desc, argpartition float16 n=1000000 desc, copy int64 n=1000000, argsort int64 n=1000000 asc, argpartition int64 n=1000000 asc, argsort int32 n=1000000 asc, argsort int16 n=1000000 desc, argpartition int16 n=1000000 desc

A runtime flag turns the comparator of the is_sorted early exit into a
function pointer, slowing ascending argsort of sorted input by ~30%.

Co-Authored-By: Claude Fable 5.1 <[email protected]>

@ngoldbaum ngoldbaum left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for following through on this @eendebakpt!

@ngoldbaum
ngoldbaum merged commit 148d78e into numpy:main Sep 21, 2026
91 checks passed
Riaz1729 pushed a commit to Riaz1729/numpy that referenced this pull request Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants