Skip to content

Commit cb0acaa

Browse files
authored
Merge pull request #182 from IntelPython/feature/add-mkl-memory
Add `MKLMemory` class to expose MKL allocated memory via Python buffer protocol
2 parents 723b6df + ce2a6bc commit cb0acaa

12 files changed

Lines changed: 1300 additions & 10 deletions

‎.github/copilot-instructions.md‎

Lines changed: 3 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -20,7 +20,7 @@ Higher-precedence file overrides; lower must not restate overridden guidance.
2020
## Contribution expectations
2121
- Keep diffs minimal; prefer atomic single-purpose commits.
2222
- Preserve public API signatures in `mkl/__init__.py` unless change is explicitly requested.
23-
- For user-visible behavior changes: update tests in `mkl/tests/test_mkl_service.py`.
23+
- For user-visible behavior changes: update tests in `mkl/tests/test_mkl_service.py`, or `mkl/tests/test_mkl_memory.py` for `MKLMemory`.
2424
- For bug fixes: add or extend regression tests in the same change.
2525
- Do not generate code without corresponding test updates when behavior changes.
2626
- Run `pre-commit run --all-files` when `.pre-commit-config.yaml` is present.
@@ -37,8 +37,8 @@ Higher-precedence file overrides; lower must not restate overridden guidance.
3737
- Build/config: `pyproject.toml`, `meson.build`
3838
- Recipe/deps: `conda-recipe/meta.yaml`, `conda-recipe/conda_build_config.yaml`
3939
- CI: `.github/workflows/*.{yml,yaml}`
40-
- API contracts: `mkl/__init__.py`, `mkl/_py_mkl_service.pyx`
41-
- Tests: `mkl/tests/test_mkl_service.py`
40+
- API contracts: `mkl/__init__.py`, `mkl/_py_mkl_service.pyx`, `mkl/_mkl_memory.pyx`
41+
- Tests: `mkl/tests/test_mkl_service.py`, `mkl/tests/test_mkl_memory.py`
4242

4343
## MKL-specific constraints
4444
- Linux runtime init path may require `RTLD_GLOBAL` preloading (`mkl/_mklinitmodule.c`).

‎AGENTS.md‎

Lines changed: 4 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -7,6 +7,7 @@ Entry point for agent context in this repo.
77
- Threading control (set/get number of threads, domain-specific threading)
88
- Version information (MKL version, build info)
99
- Memory management (peak memory usage, memory statistics)
10+
- Aligned memory allocation (`MKLMemory`, a buffer-protocol object backed by `mkl_malloc`)
1011
- Conditional Numerical Reproducibility (CNR)
1112
- Timing functions (get CPU/wall clock time)
1213
- Miscellaneous utilities (MKL_VERBOSE control, etc.)
@@ -16,6 +17,7 @@ Originally part of Intel® Distribution for Python*, now a standalone package av
1617
## Key components
1718
- **Python interface:** `mkl/__init__.py` — public API surface
1819
- **Cython wrapper:** `mkl/_py_mkl_service.pyx` — wraps MKL support functions
20+
- **Cython allocator:** `mkl/_mkl_memory.pyx` — `MKLMemory`, wraps `mkl_malloc`/`mkl_calloc`/`mkl_realloc`/`mkl_free`
1921
- **C init module:** `mkl/_mklinitmodule.c` — Linux-side MKL runtime preloading / initialization
2022
- **Helper:** `mkl/_init_helper.py` — Windows venv DLL loading helper
2123
- **Build system:** meson-python + Cython
@@ -74,11 +76,11 @@ mkl.get_version_string() # MKL version info
7476
- **API stability:** Preserve existing function signatures (widely used in ecosystem)
7577
- **Threading:** Changes to threading control must be thread-safe
7678
- **CNR:** Conditional Numerical Reproducibility flags require careful documentation
77-
- **Testing:** Add tests to `mkl/tests/test_mkl_service.py`
79+
- **Testing:** Add tests to `mkl/tests/test_mkl_service.py`, or `mkl/tests/test_mkl_memory.py` for `MKLMemory`
7880
- **Docs:** MKL support functions documented in [Intel oneMKL Developer Reference](https://www.intel.com/content/www/us/en/docs/onemkl/developer-reference-c/2025-2/support-functions.html)
7981

8082
## Code structure
81-
- **Cython layer:** `_py_mkl_service.pyx` + `_mkl_service.pxd` (C declarations)
83+
- **Cython layer:** `_py_mkl_service.pyx` and `_mkl_memory.pyx` + `_mkl_service.pxd` (C declarations)
8284
- **C init:** `_mklinitmodule.c` handles Linux preloading (`dlopen(..., RTLD_GLOBAL)`) for MKL runtime
8385
- **Windows loading helper:** `_init_helper.py` handles DLL path setup in Windows venv
8486
- **Python wrapper:** `__init__.py` imports `_py_mkl_service` (generated from `.pyx`)

‎CHANGELOG.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -10,6 +10,7 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
1010
* Enabled support of Python 3.15 [gh-243](https://github.com/IntelPython/mkl-service/pull/243)
1111
* Added support for free-threaded (GIL-disabled) CPython builds: the Cython extension is compiled with `freethreading_compatible=True` and `_mklinit` declares `Py_MOD_GIL_NOT_USED`, so importing `mkl` no longer re-enables the GIL [gh-213](https://github.com/IntelPython/mkl-service/pull/213)
1212
* Added support for new build option `ilp64` to initialize MKL with the ILP64 interface, which also resolves some build warnings [gh-184](https://github.com/IntelPython/mkl-service/pull/184)
13+
* Exposed `mkl_malloc` and related MKL calls to Python via `MKLMemory` class which supports the Python buffer protocol [gh-182](https://github.com/IntelPython/mkl-service/pull/182)
1314

1415
### Changed
1516
* Raised the minimum build-time `Cython` requirement to `3.1.0`, the first release providing the `freethreading_compatible` directive [gh-213](https://github.com/IntelPython/mkl-service/pull/213)

‎README.md‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -35,6 +35,7 @@ For more information about the usage of support functions see [Developer Referen
3535
## Building
3636

3737
A C compiler and Intel(R) oneAPI Math Kernel Library (oneMKL) are required to build mkl-service from source.
38+
The compiler must support C11 atomics (i.e., for Windows, Visual Studio 2022 17.5 or newer).
3839

3940
Executing
4041
```sh

‎meson.build‎

Lines changed: 31 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -8,6 +8,7 @@ project(
88
).stdout().strip(),
99
meson_version: '>=1.8.3',
1010
default_options: [
11+
'c_std=c11',
1112
'buildtype=release',
1213
]
1314
)
@@ -25,6 +26,20 @@ endif
2526
thread_dep = dependency('threads')
2627

2728
cc = meson.get_compiler('c')
29+
30+
atomics_args = []
31+
if cc.get_id() == 'msvc'
32+
atomics_args += '/experimental:c11atomics'
33+
endif
34+
35+
# checked to fail early if missing header
36+
if not cc.has_header('stdatomic.h', args: atomics_args)
37+
error(
38+
'mkl-service requires a C compiler supporting C11 atomics',
39+
'(i.e., for Windows, Visual Studio 2022 17.5 or newer).'
40+
)
41+
endif
42+
2843
mkl_dep = dependency('MKL', method: 'cmake',
2944
modules: ['MKL::MKL'],
3045
cmake_args: [
@@ -60,7 +75,7 @@ py.extension_module(
6075
subdir: 'mkl'
6176
)
6277

63-
# Cython extension
78+
# Cython extensions
6479
py.extension_module(
6580
'_py_mkl_service',
6681
sources: ['mkl/_py_mkl_service.pyx'],
@@ -71,6 +86,17 @@ py.extension_module(
7186
subdir: 'mkl'
7287
)
7388

89+
py.extension_module(
90+
'_mkl_memory',
91+
sources: ['mkl/_mkl_memory.pyx'],
92+
dependencies: [mkl_dep],
93+
c_args: c_args + atomics_args,
94+
link_args: rpath_link_args,
95+
install: true,
96+
subdir: 'mkl'
97+
)
98+
99+
74100
# Python sources
75101
py.install_sources(
76102
[
@@ -82,6 +108,9 @@ py.install_sources(
82108
)
83109

84110
py.install_sources(
85-
['mkl/tests/test_mkl_service.py'],
111+
[
112+
'mkl/tests/test_mkl_memory.py',
113+
'mkl/tests/test_mkl_service.py',
114+
],
86115
subdir: 'mkl/tests'
87116
)

‎mkl/AGENTS.md‎

Lines changed: 17 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -5,6 +5,7 @@ Core Python/Cython implementation: MKL support function wrappers and runtime con
55
## Structure
66
- `__init__.py` — public API, RTLD_GLOBAL context manager, module initialization
77
- `_py_mkl_service.pyx` — Cython wrappers for MKL support functions
8+
- `_mkl_memory.pyx` — `MKLMemory`, a buffer-protocol object over MKL's allocator
89
- `_mkl_service.pxd` — Cython declarations (C function signatures)
910
- `_mklinitmodule.c` — C extension for Linux-side MKL runtime preloading/init
1011
- `_init_helper.py` — Windows loading helper (DLL path setup in venv)
@@ -26,6 +27,13 @@ Core Python/Cython implementation: MKL support function wrappers and runtime con
2627
- `peak_mem_usage(memtype)` — peak memory usage stats
2728
- `mem_stat()` — memory allocation statistics
2829

30+
### Memory allocation
31+
- `MKLMemory(nbytes, alignment=64)` — aligned allocation via `mkl_malloc`; `alignment` must be a power of two
32+
- `MKLMemory(num, elem_size, alignment=64)` — zeroed allocation via `mkl_calloc`
33+
- `MKLMemory(other, alignment=other.alignment)` — copy of another allocation
34+
- `realloc(new_nbytes, refcheck=True)` — resize in place via `mkl_realloc`
35+
- `nbytes` / `__len__`, `alignment`, `tobytes()`, buffer protocol, pickling
36+
2937
### CNR (Conditional Numerical Reproducibility)
3038
- `set_num_threads_local(n)` — thread-local thread count
3139
- CNR mode control functions
@@ -39,11 +47,20 @@ Core Python/Cython implementation: MKL support function wrappers and runtime con
3947
- **API stability:** Preserve function signatures (widely used in ecosystem)
4048
- **MKL dependency:** Assumes MKL is available at runtime (conda: mkl package). Do **not** list `mkl` in `pyproject.toml` `[project].dependencies` — its PyPI wheel lacks `.dist-info`, which breaks `pip check`; on conda-forge there is no pip-visible `mkl` distribution.
4149
- **RTLD_GLOBAL preload path:** Linux preload is handled in `_mklinitmodule.c`; Windows DLL setup is in `_init_helper.py`
50+
- **`MKLMemory` mutation:** `realloc` moves the underlying block, so it must refuse while a buffer is exported, while another thread is resizing, or (unless `refcheck=False`) while the object looks referenced elsewhere. The GIL must not be released across those checks and the pointer store, mirroring NumPy's `PyArray_Resize`. The reference-count check stays NumPy's: `PyUnstable_Object_IsUniquelyReferenced` from 3.14, `Py_REFCNT > 2` before it, keyed on `PY_VERSION_HEX` and not on `Py_GIL_DISABLED`. It is a check against dangling references, not against other threads — on a free-threaded build before 3.14 it cannot be either, and resizing an allocation another thread can reach is the caller's responsibility, as it is for `numpy.ndarray.resize`.
51+
- **`MKLMemory` alignment:** `mkl_malloc`/`mkl_calloc` honor only power-of-two alignments and silently fall back to their own (64 bytes, measured) for anything else, so `_check_alignment` rejects non-powers of two — otherwise `.alignment` would report a value the allocation does not have. Powers of two are delivered exactly, up to at least 1 GiB.
52+
- **`MKLMemory` pickling:** `__reduce__` must rebuild `type(self)`, not `MKLMemory`, and carry the instance `__dict__` so a subclass survives a round trip. `_mkl_memory_from_bytes` takes the class as an optional third argument — optional so that older pickles still load, and omitted for `MKLMemory` itself so that its pickles stay loadable by older versions — and must reject anything that is not a `MKLMemory` subclass, since every pickle names that function.
53+
- **`MKLMemory` buffer export:** `__getbuffer__` hands the view to `PyBuffer_FillInfo`, which describes a flat block of unsigned bytes and answers `flags` — `format` only under `PyBUF_FORMAT`, `shape` under `PyBUF_ND`, `strides` under `PyBUF_STRIDES` — instead of filling in fields the consumer did not request. It also takes the reference on the exporter, so `__releasebuffer__` must stay a bare decrement of `exported_buffers`.
54+
- **Claim before fill:** the `atomic_fetch_add(&self.exported_buffers, 1)` comes *before* the fill, with the claim given back in an `except` clause if the fill raises (`PyBuffer_FillInfo` is declared `except -1`). Claiming afterwards leaves a window in which a concurrent `realloc` frees the block the view was already handed, and the consumer keeps that view — every array over the allocation holds it for as long as the array lives, so the cost is a durably dangling array rather than one bad read. Reproduced with the window widened by a 5 ms sleep on 3.13t: the resize went through and ASan reported `heap-use-after-free` in `array_tobytes`; with the claim first the same resize is refused. No test can observe the ordering, so it has to be kept on purpose. It narrows rather than closes the race — a `realloc` already past its own count check can still free under a fill — which only mutual exclusion would fix.
55+
- **Backing a NumPy array:** `np.asarray(mem)` and `np.frombuffer(mem, dtype=...)` keep a `memoryview` as `.base` and hold the export for the array's whole lifetime, so `realloc` is refused with `BufferError` until the array goes away. `np.ndarray(shape, buffer=mem)` releases the `Py_buffer` and keeps only an object reference, so only the reference check stands in the way and `refcheck=False` leaves the array dangling — `bytearray` behaves the same there, so it is NumPy's property, not this object's. The array cannot resize the allocation either: it does not own its data, which `PyArray_Resize` refuses ahead of its own reference check, so neither `ndarray.resize(..., refcheck=False)` nor a C caller invoking `PyArray_Resize` directly gets past it (both measured). What does drop the export a live array depends on is `arr.base.release()`, which is caller error the same way it is for any exporter.
56+
- **Reading another `MKLMemory`'s block:** code that reads someone else's allocation with the GIL released must claim a buffer on it (`atomic_fetch_add(&other.exported_buffers, 1)` in a `try`/`finally`, as the copy constructor does) *before* reading its size, so that a concurrent `realloc` is refused rather than freeing the block mid-read or shrinking it under a size that was already read.
4257

4358
## Cython details
4459
- `_py_mkl_service.pyx` → generates `_py_mkl_service` extension module
60+
- `_mkl_memory.pyx` → generates `_mkl_memory` extension module
4561
- `.pxd` file declares external C functions from MKL headers
4662
- Cython build requires MKL headers (`mkl-devel`)
63+
- `_mkl_memory.pyx` uses C11 atomics (`<stdatomic.h>`); `meson.build` scopes MSVC's `/experimental:c11atomics` to that one target
4764

4865
## C init module
4966
- `_mklinitmodule.c` → `_mklinit` extension

‎mkl/__init__.py‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -57,6 +57,7 @@ def __exit__(self, *args):
5757

5858
del RTLD_for_MKL
5959

60+
from ._mkl_memory import MKLMemory
6061
from ._py_mkl_service import (
6162
cbwr_get,
6263
cbwr_get_auto_branch,
@@ -121,6 +122,7 @@ def __exit__(self, *args):
121122
"mem_stat",
122123
"peak_mem_usage",
123124
"set_memory_limit",
125+
"MKLMemory",
124126
"cbwr_set",
125127
"cbwr_get",
126128
"cbwr_get_auto_branch",

0 commit comments

Comments
 (0)