Small examples that exercise PyCuTe end-to-end. Run commands below from the
repository root so that the examples package and local pycute sources are on
the import path. A virtual environment is recommended; see
Installation.
Regenerates the SVG figures embedded in the top-level README:
row-major, column-major, and blocked layouts; an SM80 16×8 thread-value
layout; and two shared-memory layouts with F2 (XOR) swizzled strides.
Install the visualization dependencies, then run:
pip install -e ".[viz]"
python3 -m examples.readme_figuresThe script writes SVGs to docs/images/ regardless of the
current working directory. See
Visualization for the drawing APIs.
Rearranges modes by naming them: einfold("in_modes->out_modes", value) is a
generalized transpose that permutes, groups, ungroups, repeats and drops the
modes of a Layout, a Tensor or an HTuple. Each side of the expression
denotes a profile whose leaves are one-character names — as in einsum's
subscripts, separators between modes are implied — so a group descends into a
nested source mode on the left and builds a new one on the right. Top-level
modes the input leaves unnamed are appended unchanged, so an expression only
names the modes it rearranges.
Layouts produce Layouts and Tensors produce Tensors over the source's
accessor, so no data is copied. Anything else is matched against its
profile and
rebuilt as tuples, which asks nothing of its leaves — so a Shape or a Stride
folds exactly as the Layout it was read from.
from pycute import Layout, make_tensor, shape
from examples.einfold import einfold
A = make_tensor(Layout(((12, 4), 42, 5, 7)))
shape(einfold("ijkm -> imkj", A)) # ((12, 4), 7, 5, 42) -- permute
shape(einfold("(ab)cde -> c(ade)b", A)) # (42, (12, 5, 7), 4) -- split and regroup
shape(einfold("ij -> (ij)", A)) # (((12, 4), 42), 5, 7) -- 5, 7 pass through
shape(einfold("ijk -> ik", A)) # ((12, 4), 5, 7) -- drop mode j
einfold("ijkm -> imkj", shape(A)) # folds the Shape on its ownThe companion einfold_test.py is one series of such examples, each checked as
a Layout, as a Tensor and as the source's Shape and Stride:
pytest examples/einfold_test.pyImplements a minimal binary einsum by folding a tensor contraction into the
reference batched GEMM,
pycute.alg.ref.gemm. Given
einsum("a_modes,b_modes->c_modes", A, B, C), labels are classified as row
(M), column (N), reduction (K), or batch (L) modes, then viewed as the
canonical A:(M,K,L), B:(N,K,L), and C:(M,N,L) layouts.
C is caller-provided and the result accumulates into it. Only explicit
subscripts (->) are supported; labels may appear at most once per operand,
there is no broadcasting, and each label must occur in at least two operands.
from pycute import Layout, make_tensor
from examples.einsum import einsum
A = make_tensor(Layout((3, 5))) # m,k
B = make_tensor(Layout((4, 5))) # n,k
C = make_tensor(Layout((3, 4))) # m,n; initially zero
einsum("mk,nk->mn", A, B, C)The companion einsum_test.py uses a dependency-free brute-force oracle:
pytest examples/einsum_test.pyThe im2col transformation as a Layout, so an N-D convolution runs as a GEMM
over the activation in place -- what CUTLASS calls implicit GEMM. The
activation's spatial modes are composed twice, once with a position tiler strided
by the traversal stride and once with a tap tiler strided by the dilation, giving
a rank-2 ((N,(Z,P,Q)), ((T,R,S),C)) view that
pycute.alg.ref.gemm consumes unmodified.
Three views share one construction, differing only in what the layout maps into:
im2col to offsets, im2col_coord to activation coordinates (what predication
and a TMA descriptor need), and im2col_padded to coordinates read through a
bounds-checking accessor, which is how a padded convolution reaches an
unmodified GEMM.
from pycute import Layout, make_tensor
from examples.im2col import im2col
act = make_tensor(Layout((1, 4, 4, 1), (16, 4, 1, 1))) # (N,H,W,C)
A = im2col(act, (2, 2)) # ((1,(3,3)), ((2,2),1))The companion im2col_test.py checks it against a direct N-D
cross-correlation, over 1-D through 3-D, traversal stride, dilation, symmetric
and asymmetric padding, dgrad coordinates and CUTLASS example 59's layout:
pytest examples/im2col_test.pyWalkthrough of the COPY algorithm (Whitepaper §2.6.1): applications that are
all copy(src, dst) with different layouts (memcpy, gather/scatter, broadcast,
transpose, …), then the layout analysis pycute.alg.copy uses to reshape that
loop — common domain, nullspace, and alignment — against the element-at-a-time
pycute.alg.ref.copy.
pip install -e ".[viz]" # optional inline SVG layout figures
jupyter notebook examples/algorithms/copy.ipynbUnlike the scripts above, the notebook does not need to run from the repository
root: when pycute is not importable it walks up from the kernel's working
directory to the checkout and puts that on sys.path, so a bare
jupyter notebook works without installing anything.
Unit coverage for both loops lives in test/test_alg_copy.py, which also runs
every copy the notebook shows against the reference, pins its Stage 3 table,
and checks that its stored source listings are current.
Walkthrough of the GEMM algorithm (Whitepaper §2.6.2), in which every
application is one call to pycute.alg.ref.gemm:
the BLAS transpose variants as a stride choice rather than an algorithm choice,
tensor folding and the einsum applications reviewed above, and then CONV.
The CONV half is a tutorial on implicit GEMM, built on
im2col.py above: N-D stencils, traversal stride, dilation,
padding through a bounds-checking accessor, and dgrad. Ten convolutions run as
one gemm call each, checked against a direct cross-correlation oracle, with no
data movement anywhere.
pip install -e ".[viz]" # optional inline SVG layout figures
jupyter notebook examples/algorithms/gemm.ipynbLike copy.ipynb it runs from any directory in the checkout, walking up to the
repository root when examples is not importable. test/test_alg_gemm.py
checks that its stored source listings are current.
Keep each example self-contained, import from pycute (and pycute.util for
visualization), and add a short section here describing how to run it and what
it produces.