Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Intra-Slice Chain

The Intra-Slice Chain performs elementwise, binary, and intra-slice reduce operations on each slice’s data independently. It handles post-contraction processing such as activation and normalization. For example, computing \(\operatorname{sigmoid}(XW + b)\) runs \(XW\) on the Contraction Engine, and the addition plus sigmoid activation in the Intra-Slice Chain.

Interface

The example below applies a fixed-point bias, runs sigmoid on the float path (narrowing to 4-way and widening back), and finishes with a ReLU clip. It computes \(output[a, b] = \max(\operatorname{sigmoid}(input[a, b] + 100), 0)\).

#![allow(unused)]
fn main() {
#![feature(adt_const_params)]
extern crate furiosa_opt_std;
use furiosa_opt_std::prelude::*;
axes![A = 512, B = 2];

fn staged_pipeline<'l, const T: Tu>(
    input: CollectTensor<'l, T, i32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]>,
) -> VectorFinalTensor<'l, T, i32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]> {
    input
    .vector_init()
    .vector_intra_slice_tag(TagMode::Zero)
    // input + 100
    .vector_fxp(FxpBinaryOp::AddFxp, 100)
    // i32 → f32 (fixed-point, int_width = 31)
    .vector_fxp_to_fp(31)
    // Way8 → Way4 for the float path
    .vector_narrow_trim::<m![A % 2 # 4]>()
    // sigmoid(input + 100)
    .vector_fp_unary(FpUnaryOp::Sigmoid)
    // Way4 → Way8
    .vector_widen_pad::<m![A % 2 # 8]>()
    // f32 → i32
    .vector_fp_to_fxp(31)
    // max(sigmoid(input + 100), 0)
    .vector_clip(ClipBinaryOpI32::Max, 0)
    .vector_final()
}

let mut ctx = Context::acquire();

let c: CollectTensor<'_, _, i32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]> = CollectTensor::new(&mut ctx.main, Tensor::zero());
let _o = staged_pipeline(c);
}

Pipeline

The chain starts (as shown above) with vector_intra_slice_tag() (right after vector_init() or on the inter-slice reducer’s output), or with vector_intra_slice_unzip() directly after vector_init() for Pair Mode (see below).

    #[primitive(VectorInitTensor::vector_intra_slice_tag)]
    pub fn vector_intra_slice_tag(
        self,
        branch: TagMode<D>,
    ) -> VectorBranchTensor<'l, T, D, Chip, Cluster, Slice, Time, Packet, Fresh, { VeOrder::IntraFirst }> {

    #[primitive(VectorInitTensor::vector_intra_slice_unzip)]
    pub fn vector_intra_slice_unzip<I: AxisName, TileTime: M, SplitTime: M>(
        self,
    ) -> VectorTensorPair<'l, T, D, stage::Tag, Chip, Cluster, Slice, SplitTime, Packet> {

After entry, the chain steps through the pipeline stages below in a fixed order; software chains the relevant ones and skips the rest, as the example skips Logic, FpDiv, and Filter. Each row lists the stage’s position in the chain (#), its name (Stage), the API method that triggers it (Method), the way it runs in (Way, either 8 or 4 elements per cycle), and whether it accepts an operand (Operand). The type system enforces every stage transition at compile time, so methods become callable only after the preceding chain reaches a compatible state. Per-stage detail is in Stages below.

#StageMethodWayOperand→ Inter-Slice Reducer
1Entryvector_intra_slice_tag()Way8
2Logicvector_logic()Way8yesyes
3Fxpvector_fxp()Way8yesyes
4FxpToFpvector_fxp_to_fp()Way8yes
5Narrowvector_narrow_split() / vector_narrow_trim()Way8 → Way4
6Floatvector_fp_unary/binary/ternary()Way4yes
7IntraSliceReducevector_intra_slice_reduce()Way4
8FpDivvector_fp_div()Way4yes
9Widenvector_widen_concat() / vector_widen_pad()Way4 → Way8yes
10FpToFxpvector_fp_to_fxp()Way8yes
11Clipvector_clip()Way8yesyes
12Filtervector_filter()Way8

vector_reinterpret() is deliberately absent from the table: it occupies no stage and can sit between any two of them (see Reinterpret).

Stages run either 8-way (8 elements per cycle) or 4-way (4 elements per cycle). The floating-point cluster runs 4-way to amortize its half-throughput ALUs against the rest of the chain. A chain that uses the float path therefore enters 8-way, calls Narrow (vector_narrow_split or vector_narrow_trim) before the float stages, and calls Widen (vector_widen_concat or vector_widen_pad) afterward to return to 8-way. The example does exactly this: vector_narrow_trim then vector_fp_unary(Sigmoid) then vector_widen_pad.

The chain exits via vector_final() (as in the example) or vector_inter_slice_reduce() (handing off to the Inter-Slice Reducer). Both exits require 8-way, so any active 4-way stage must pass through Widen first.

Operands

Binary and ternary ops have two slot kinds: a stream (the running tensor, i.e., the self of the method chain, fixed by the chain) and one or two operands (the extra inputs). Each operand comes from one of three sources.

SourceExampleDescription
Constant100, 2.5f32Scalar broadcast to all elements
VRF tensor&vrf_tensorPre-loaded via .to_vrf() before entering the Vector Engine
StashStashSnapshot of an earlier chain step, described below

Ternary ops (FmaF) take a pair (operand0, operand1).

A VRF operand is read per slice by an indexer, which puts three rules on it.

First, its Chip / Cluster / Slice are the stream’s own: at slice s the op reads what slice s holds, so an operand under a different partition is a compile error instead of a silent read of another slice’s data. VrfTensor::reshape restates one under the stream’s partition when the two describe the same slices, e.g. an operand replicated under an anonymous m![256] feeding a stream partitioned by a named axis.

Second, one access (the Packet: 8 lanes in 8-way, 4 in 4-way) is one indexer step, and the indexer offers just two read options. Broadcast feeds every lane from one address; contiguous reads Packet consecutive Element cells. So the packet axes must be the operand’s innermost, contiguous ones, or absent from it entirely; a packet that walks an outer axis of the operand, or mixes broadcast and live lanes inside one access, has no encoding and is rejected.

Third, the two directions of an axis mismatch are not the same. An axis the operand has and the stream does not is refused, with the stream leaves operand axes unmatched, because the cells it indexes would never be read. An axis the stream has and the operand does not is a broadcast, which is how one operand feeds every row of the stream.

The same op method picks up different sources by the argument type:

.vector_fxp(FxpBinaryOp::AddFxp, 100)         // operand from constant
.vector_fxp(FxpBinaryOp::MulInt, &vrf)        // operand from VRF tensor
.vector_clip(ClipBinaryOpI32::Max, Stash)     // operand from stash (set earlier)

The Stash source comes from vector_stash(), which snapshots the running tensor so a later binary or ternary op can read it back as the Stash operand. The typical use is a residual or skip-connection like max(f(x), x), where the original x must survive across intermediate stages. Although stash is presented as a separate operand source, it is backed by the VRF: part of the current running tensor is stored in the VRF and read back from there. Whenever a Tensor Unit invocation uses stash, the compiler conservatively reserves 1,024 B of VRF capacity per slice for it, regardless of the actual amount of data stashed. If that invocation also reads a pre-loaded VRF tensor as an RHS operand, size the operand with this reservation in mind: stash and the RHS share the 8 KiB VRF, so at most 7 KiB per slice remains for the RHS. Call vector_stash() at any Stashable stage (Branch, Logic, Fxp, Narrow, Fp, FpDiv, Clip); the snapshot stays live until the Tensor Unit invocation ends and feeds any later binary or ternary call that takes Stash. The slot is single-use (a second vector_stash() is a compile-time error) and typed, so an f32 stash only feeds f32 ops: a conversion or reinterpret between the write and the read has to be undone first. The stash is also read-once: it feeds exactly one later op (reading it moves the slot past Occupied, so a second Stash read is a compile-time error). A value that must be read more than once is not a stash: put it in a read-many VRF, passed as &vrf. The mapping follows the running tensor, so a stash taken before Narrow is still usable after Widen. Those seven stages are the seven chain steps that own a VRF write port, which is why Widen is not one of them: the operand register cannot be written from the widen’s output at all. A value computed in the 4-way region therefore reaches the stash by widening first and stashing at a Clip op, which puts both the write and the read on the 8-way side. Stash is unavailable in Pair Mode, and the IntraFirst transition to the inter-slice reducer drops it, so anything stashed before vector_inter_slice_reduce() is gone afterward.

An argument mode then picks which slots hold the stream versus the operands so the same op can compute, e.g., stream + operand or operand - stream. For example, BinaryArgMode::Mode10 swaps the slots so SubFxp computes operand - stream:

.vector_fxp_with_mode(FxpBinaryOp::SubFxp, BinaryArgMode::Mode10, 7)  // computes 7 - stream

BinaryArgMode picks which two of a binary op’s slots are stream vs operand (unary ops have no mode and always run as op(stream)):

BinaryArgModeSlotsComputation
Mode00stream / streamop(stream, stream)
Mode01stream / operandop(stream, operand) (default)
Mode10operand / streamop(operand, stream)
Mode11operand / operandop(operand, operand)

TernaryArgMode does the same for ternary ops:

TernaryArgModeSlotsComputation
Mode012stream / operand0 / operand1op(stream, operand0, operand1) (default)
Mode002stream / stream / operand1op(stream, stream, operand1)
Mode102operand0 / stream / operand1op(operand0, stream, operand1)
Mode112operand0 / operand0 / operand1op(operand0, operand0, operand1)
Mode020stream / operand1 / streamop(stream, operand1, stream)
Mode021stream / operand1 / operand0op(stream, operand1, operand0)
Mode120operand0 / operand1 / streamop(operand0, operand1, stream)

Pair Mode

Pair mode runs the chain on a tensor whose elements split into two interleaved groups, so an op can relate the two groups (e.g., pair-wise add, asymmetric scale). Entry is vector_intra_slice_unzip() directly after vector_init(), applied to a collected tensor that carries a 2-way grouping axis. Starting the chain with vector_intra_slice_unzip() precludes the Filter stage downstream. Under the hood, vector_intra_slice_unzip() derives each element’s GroupId from its position along the 2-way grouping axis, driving the branch unit’s count augmenter rather than any TagMode. The tag the chain was started with is untouched, which is why this is the supported route to a group split and TagMode::AxisToggle is not.

The flow has four steps:

  1. vector_intra_slice_unzip() splits the input into two parallel streams (group 0 and group 1).
  2. The chain runs through stages with both groups in lock-step (the paired phase).
  3. A _zip op fuses the two streams back into one (the merged phase).
  4. The merged stream continues to vector_final() like a normal chain.

Stages during the paired phase fall into two flavors:

  • Common stages (vector_fxp_to_fp, vector_narrow_split, vector_widen_concat, vector_fp_to_fxp) act on both groups uniformly. vector_narrow_trim and vector_widen_pad are not available on pairs; use the _split / _concat variants instead.
  • Per-group ops take one argument per group:
    • Binary and ternary (vector_fxp, vector_fp_binary, vector_fp_ternary, vector_clip, etc.) accept () on a side to skip it, or different operands on each side.
    • Unary (vector_fp_unary) is the exception: it takes flags (op, group0_apply, group1_apply) and false skips that group.

Pair mode reinterprets BinaryArgMode depending on the op: per-group ops (vector_fxp_with_mode, vector_fp_binary_with_mode, etc.) apply the mode inside each group independently (0 is that group’s stream, 1 is that group’s operand), while _zip ops (vector_fxp_zip_with_mode, etc.) take the two slots as the two grouped streams (0 is Group 0’s stream, 1 is Group 1’s stream):

_zip BinaryArgModeSlotsComputation
Mode00group0 / group0op(group0, group0)
Mode01group0 / group1op(group0, group1) (default)
Mode10group1 / group0op(group1, group0)
Mode11group1 / group1op(group1, group1)

Pair-mode constraints:

  • stash() and filter() are unavailable throughout pair mode (both paired and merged phases).
  • vector_intra_slice_reduce() is unavailable throughout pair mode as well, before and after the _zip: the reducer takes no group condition.
  • Before _zip (the paired phase), the chain cannot transition to the inter-slice reducer, since vector_inter_slice_reduce() is not available on per-group tensors. After _zip (the merged phase), the result is Commitable again and can call vector_inter_slice_reduce() if the current stage supports the transition.
  • ALU usage is shared across the two groups: an ALU used in either group counts as consumed for both.

Stages

Within a stage, each ALU runs at most once per Tensor Unit invocation. This matters mainly in Logic, Fxp, Fp, and Clip, where multiple operators share a stage-local ALU pool. For example, tanh(sqrt(x)) cannot fit in a single Tensor Unit invocation because both tanh and sqrt consume the FpFpu ALU.

Tag

The Tag stage is the chain’s entry point and assigns each 32-bit element in a flit a 4-bit Tag (0-15), which later stages use to apply conditional operations. Bit 3 (the MSB) is GroupId, used by Filter and pair mode to split elements into Group 0 / Group 1. Bits 0..2 are general-purpose flag bits filled by comparison results.

The TagMode selects how the 4 bits are computed for each element. Only Zero and Comparison are available today. AxisToggle is declared but not compilable, as the last column records; ValidCount and Vrf are withheld from the enum entirely until they run on both paths.

TagModeHow each tag bit is filledStatus
ZeroAll four bits are 0. Every element has tag = 0.Works
Comparison([cmp0, cmp1, cmp2, cmp3])For each element x, bit i = 1 iff cmp_i(x) holds. Each cmp_i is a Cmp<D> carrying the boundary it compares against, typed by the stream’s scalar so the boundary is checked against it: Equal, Less, Greater, LessUnsigned, GreaterUnsigned, True, False. The *Unsigned pair compares raw bit patterns.Works
AxisToggle { axis }Bit 3 (GroupId) = axis_index % 2 along axis. Bits 0..2 stay 0.Not compilable. Runs on the host, but rejected with a kernel compile error when lowered. vector_intra_slice_unzip() is the supported way to get this grouping.

For example, with i32 data and Comparison([Cmp::Less(0), Cmp::Equal(5), Cmp::Greater(100), Cmp::True]), an element x = 7 yields bits 0/0/0/1 (LSB first), so its tag is 0b1000 = 8.

Once tags are assigned, later binary and ternary ops can condition an operand on the tag, so different elements see different values. The rule is one sentence: with no condition write the operand bare, and with a condition name the hardware slot it drives, through Branched. ve_elementwise_branched in furiosa-opt-examples is a compiled kernel doing exactly that – a guarded immediate with a trailing unconditional rf read – and ve_branch_bit_order adds three immediates and a guarded stash read. Both are read here rather than restated, so an API change breaks them instead of leaving prose behind.

A TagGuard is TagGuard::all(), TagGuard::group(id) for one group, or matches/not_matches of a [BitReq; 4] pattern written LSB first, one BitReq per tag bit: One where the bit must be 1, Zero where it must be 0, Ignore where it does not matter. The requirement is on the bit’s value, whatever the tag mode filled it from.

A pass has four operand slots: three immediate registers and the register-file port, which reads either a VRF register or the stash. A conditional operand says which kind of slot it drives — Branched::imm for an immediate, Branched::rf for the port — and gets back a builder offering only what may still follow. imm may be called three times and rf once, last; a fourth imm, or an imm after rf, does not compile.

The kernel does not pick which immediate register: imm takes the next one, and which physical register a value ends up in is command generation’s choice. What is preserved is call order — the slots reach the hardware in the order they were filled, and that order is match priority, so an element takes the first slot whose guard it satisfies and an unconditional slot is the pass’s else with nothing after it.

Two guards are rejected outright, because a slot carrying one is filled and inert while a pass has only four: TagGuard::not_matches([Ignore, Ignore, Ignore, Ignore]), which no execution id satisfies, and any slot placed after an unconditional one. TagGuard::matches([Ignore, Ignore, Ignore, Ignore]) is accepted and stored as TagGuard::all(), the same predicate spelled shorter.

Helpers that build a layout and hand it back are ordinary functions, so a kernel library can wrap whatever combinations it uses often. Two things such a helper cannot do, both of which the operand module documents in full: it cannot introduce a new operand type (an operand carries a claim about what it does to the stash, so the traits that say so are closed), and it cannot fill a slot conditionally, because how many slots are spoken for is part of the builder’s type and the two arms of an if would disagree. Vary the payload or the guard instead of the slot count, or branch around the whole op.

Logic Cluster

The Logic Cluster performs bitwise operations on i32 or f32 (bit-level). It runs 8-way.

The stage exposes five ALU classes (LogicAnd, LogicOr, LogicXor, LogicLshift, LogicRshift), each runnable once per Tensor Unit invocation. Operators sharing the same class cannot fuse into one invocation.

i32 operations:

OpALUNote
BitAndLogicAndbitwise and
BitOrLogicOrbitwise or
BitXorLogicXorbitwise xor
LeftShiftLogicLshiftlogical left shift
LogicRightShiftLogicRshiftlogical right shift
ArithRightShiftLogicRshiftarithmetic right shift

f32 operations:

OpALUNote
BitAndLogicAndbitwise and on fp bit patterns
BitOrLogicOrbitwise or on fp bit patterns
BitXorLogicXorbitwise xor on fp bit patterns

Fxp Cluster

The Fxp Cluster performs integer and fixed-point arithmetic on i32. It runs 8-way.

The stage exposes four ALU classes (FxpAdd, FxpLshift, FxpMul, FxpRshift), each runnable once per Tensor Unit invocation. Operators sharing the same class cannot fuse into one invocation.

OpALUNote
AddFxpFxpAddwrapping add
AddFxpSatFxpAddsaturating add
SubFxpFxpAddwrapping subtract
SubFxpSatFxpAddsaturating subtract
LeftShiftFxpLshiftlogical left shift
LeftShiftSatFxpLshiftsaturating left shift
MulFxpFxpMulfixed-point multiply
MulIntFxpMulinteger multiply
LogicRightShiftFxpRshiftlogical right shift
ArithRightShiftFxpRshiftarithmetic right shift
ArithRightShiftRoundFxpRshiftarithmetic right shift with rounding

The single-ALU rule rejects, for example, two ops that both target FxpAdd:

// PANICS: "FxpAdd is already in use"
input
    .vector_init()
    .vector_intra_slice_tag(TagMode::Zero)
    .vector_fxp(FxpBinaryOp::AddFxp, 10)    // uses FxpAdd
    .vector_fxp(FxpBinaryOp::MulInt, 2)     // uses FxpMul ✓
    .vector_fxp(FxpBinaryOp::SubFxp, 5)     // uses FxpAdd again ✗
    .vector_final()

FxpToFp Conversion

The FxpToFp Conversion stage converts i32 to f32. The int_width parameter specifies the integer bit width for the conversion; int_width = 31 is the standard i32f32 conversion.

MethodEffect
vector_fxp_to_fp(int_width)convert i32 stream to f32

Narrow

The Narrow stage switches 8-way to 4-way. An 8-way packet carries 8 active elements (Packet = m![... # 8]) and a 4-way packet carries 4 (Packet = m![... # 4]). Narrowing halves throughput on the float and reduce path, so the same logical tensor shape takes twice as many packets or Tensor Unit invocations.

MethodUse WhenEffect
vector_narrow_split()both halves contain real datasplit one 8-way flit into a front-4 and back-4 packet, updating Time and Packet
vector_narrow_trim()back 4 elements are already padding or irrelevantkeep only the front 4 elements

Shape semantics:

#![allow(unused)]
fn main() {
#![feature(adt_const_params)]
extern crate furiosa_opt_std;
use furiosa_opt_std::prelude::*;
axes![A = 512, B = 2, S = 64];

fn vector_narrow_split_semantics<'l, const T: Tu>(
    input: VectorBranchTensor<'l, T, i32, m![1], m![B], m![S / 4 # 256], m![S % 4], m![A % 8], Fresh, { stage::VeOrder::IntraFirst }>,
) -> VectorNarrowTensor<'l, T, i32, m![1], m![B], m![S / 4 # 256], m![S % 4, A / 4 % 2], m![A % 4], Fresh, { stage::VeOrder::IntraFirst }>
{
    input.vector_narrow_split::<m![S % 4, A / 4 % 2], m![A % 4]>()
    // shape semantics: [T], [P] -> [T, P / 2], [P % 4]
}

fn vector_narrow_trim_semantics<'l, const T: Tu>(
    input: VectorBranchTensor<'l, T, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8], Fresh, { stage::VeOrder::IntraFirst }>,
) -> VectorNarrowTensor<'l, T, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 4], Fresh, { stage::VeOrder::IntraFirst }>
{
    input.vector_narrow_trim::<m![A % 2 # 4]>()
    // shape semantics: [T], [P] -> [T], [P = 4]
}

let mut ctx = Context::acquire();

let i: VectorBranchTensor<'_, _, i32, m![1], m![B], m![S / 4 # 256], m![S % 4], m![A % 8], Fresh, { stage::VeOrder::IntraFirst }> = VectorBranchTensor::new(&mut ctx.main, Tensor::zero(), TagMode::Zero);
let _o = vector_narrow_split_semantics(i);

let i: VectorBranchTensor<'_, _, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8], Fresh, { stage::VeOrder::IntraFirst }> = VectorBranchTensor::new(&mut ctx.main, Tensor::zero(), TagMode::Zero);
let _o = vector_narrow_trim_semantics(i);
}

Float Cluster

The Float Cluster provides unary, binary, and ternary floating-point operations on f32. It runs 4-way, so the input must already have passed through Narrow.

It exposes five independent ALUs (FpFma, FpFpu, FpExp, FpMul0, FpMul1), each runnable once per Tensor Unit invocation. This stage is where ALU planning matters most.

Unary ops:

OpALUNote
ExpFpExpexponential
NegExpFpExpnegative exponential
SqrtFpFpusquare root
TanhFpFpuhyperbolic tangent
SigmoidFpFpusigmoid
ErfFpFpuerror function
LogFpFpunatural logarithm
SinFpFpusine
CosFpFpucosine

Binary ops:

OpALUNote
AddFFpFmafloating-point add
SubFFpFmafloating-point subtract
MulF(FpMulAlu::Mul0)FpMul0multiply
MulF(FpMulAlu::Mul1)FpMul1multiply
MulF(FpMulAlu::Fma)FpFmamultiply
DivFFpFpudivision inside Fp stage

Ternary ops:

OpALUNote
FmaFFpFmafused multiply-add

For example, to compute exp(sqrt(((x + 1) * 2) * 3)):

  • x1 = x + 1 via FpFma (FpBinaryOp::AddF)
  • x2 = x1 * 2 via FpMul0 (FpBinaryOp::MulF(FpMulAlu::Mul0))
  • x3 = x2 * 3 via FpMul1 (FpBinaryOp::MulF(FpMulAlu::Mul1))
  • x4 = sqrt(x3) via FpFpu (FpUnaryOp::Sqrt)
  • x5 = exp(x4) via FpExp (FpUnaryOp::Exp)

IntraSliceReduce

The IntraSliceReduce stage reduces axes within a single slice. It runs 4-way. The stage uses a dedicated accumulator-tree ALU, so no user-selectable ALU is exposed.

Data TypeSupported Ops
i32AddSat, Max, Min
f32Add, Max, Min

See Intra-Slice Reduce for details.

FpDiv

The FpDiv stage performs floating-point division. It runs 4-way. The stage uses a dedicated floating-point divider, so no user-selectable ALU is exposed.

OpNote
FpDivBinaryOp::DivFdedicated floating-point division

Widen

The Widen stage transitions from 4-way back to 8-way. Later stages (FpToFxp, Clip, Filter, Output) then see 8-element packets again.

MethodUse WhenEffect
vector_widen_concat()reversing a prior vector_narrow_split()merge two 4-way packets back into one 8-way flit
vector_widen_pad()reversing a prior vector_narrow_trim()pad a 4-way packet back to 8 elements with invalid fillers

Shape semantics:

#![allow(unused)]
fn main() {
#![feature(adt_const_params)]
extern crate furiosa_opt_std;
use furiosa_opt_std::prelude::*;
axes![A = 512, B = 2, S = 64, R = 8];

fn vector_widen_concat_semantics<'l, const T: Tu>(
    input: VectorIntraSliceReduceTensor<'l, T, i32, m![1], m![B], m![S / 4 # 256], m![A / 4 % 2], m![A % 4], Fresh, { stage::VeOrder::IntraFirst }>,
) -> VectorWidenTensor<'l, T, i32, m![1], m![B], m![S / 4 # 256], m![1], m![A % 8], Fresh, { stage::VeOrder::IntraFirst }>
{
    input.vector_widen_concat::<m![1], m![A % 8]>()
    // shape semantics: [T, P / 2], [P % 4] -> [T], [P]
}

fn vector_widen_pad_semantics<'l, const T: Tu>(
    input: VectorFpTensor<'l, T, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 4], Fresh, { stage::VeOrder::IntraFirst }>,
) -> VectorWidenTensor<'l, T, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8], Fresh, { stage::VeOrder::IntraFirst }>
{
    input.vector_widen_pad::<m![A % 2 # 8]>()
    // shape semantics: [T], [P] -> [T], [P # 8]
}

let mut ctx = Context::acquire();

let i: VectorBranchTensor<'_, _, i32, m![1], m![B], m![S / 4 # 256], m![R, A / 4 % 2], m![A % 4 # 8], Fresh, { stage::VeOrder::IntraFirst }> = VectorBranchTensor::new(&mut ctx.main, Tensor::zero(), TagMode::Zero);
let i = i
    .vector_narrow_trim::<m![A % 4]>()
    .vector_intra_slice_reduce::<R, m![A / 4 % 2], m![A % 4]>(IntraSliceReduceOpI32::AddSat);

let _o = vector_widen_concat_semantics(i);

let i: VectorBranchTensor<'_, _, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8], Fresh, { stage::VeOrder::IntraFirst }> = VectorBranchTensor::new(&mut ctx.main, Tensor::zero(), TagMode::Zero);
let i = i.vector_narrow_trim::<m![A % 2 # 4]>().vector_fp_unary(FpUnaryOp::Exp);
let _o = vector_widen_pad_semantics(i);
}

FpToFxp Conversion

The FpToFxp Conversion stage converts f32 back to i32. The int_width parameter specifies the integer bit width.

MethodEffect
vector_fp_to_fxp(int_width)convert f32 stream back to i32

Clip Cluster

The Clip Cluster performs clamping and comparison operations. It runs 8-way.

The stage exposes three ALU classes (ClipAdd, ClipMax, ClipMin), each runnable once per Tensor Unit invocation.

i32 operations:

OpALUNote
MinClipMinminimum
MaxClipMaxmaximum
AbsMinClipMinabsolute minimum
AbsMaxClipMaxabsolute maximum
AddFxpClipAddwrapping add
AddFxpSatClipAddsaturating add

f32 operations:

OpALUNote
MinClipMinminimum
MaxClipMaxmaximum
AbsMinClipMinabsolute minimum
AbsMaxClipMaxabsolute maximum
AddClipAddfloating-point add

Filter

The Filter stage applies an execution mask derived from a TagGuard (typically matching on the GroupId MSB of each element’s Tag) to filter output flits. Available 8-way and Standalone context only. The source impl lives on VectorTensor for any stage with CanTransitionTo<Filter>, which covers every intra-slice stage and InterSliceReduce.

Output

The Output stage exits the Vector Engine pipeline. The result can continue to the Cast Engine, Transpose Engine, or Commit Engine.

Reinterpret

vector_reinterpret::<D2>() rereads the stream as the other 32-bit scalar with the bits untouched, so 1.0f32 becomes 0x3f80_0000 and not 1. Every element is 32 bits and every cluster takes its operand type from the op it runs, so a reinterpret emits no instruction, claims no ALU, and holds both the current stage and the way. It can therefore appear anywhere in the chain, any number of times, including between two ops of the same cluster.

One possible use for a reinterpret is applying a bitwise operation to floating-point values, as in Bitwise abs on f32. For the value-preserving conversions use FxpToFp / FpToFxp instead.

MethodEffect
vector_reinterpret::<i32>()read the f32 stream’s bits as i32
vector_reinterpret::<f32>()read the i32 stream’s bits as f32

A reinterpret also decides what a following vector_stash() writes, since the stash takes the scalar the stream carries at the write. So a value reaches a reader of the other scalar by reinterpreting first and stashing after; a read at the other scalar does not compile.

Examples

i32 Pipeline

A minimal i32 chain that adds a constant after branching. The Fxp stage runs 8-way, so no narrow or widen is needed.

#![allow(unused)]
fn main() {
#![feature(adt_const_params)]
extern crate furiosa_opt_std;
use furiosa_opt_std::prelude::*;
axes![A = 2048, B = 2];

fn add_constant<'l, const T: Tu>(
    input: CollectTensor<'l, T, i32, m![1], m![B], m![A / 8], m![1], m![A % 8]>,
) -> VectorFinalTensor<'l, T, i32, m![1], m![B], m![A / 8], m![1], m![A % 8]> {
    input
        .vector_init()
        .vector_intra_slice_tag(TagMode::Zero)
        .vector_fxp(FxpBinaryOp::AddFxp, 100)
        .vector_final()
}

let mut ctx = Context::acquire();

let i: CollectTensor<'_, _, i32, m![1], m![B], m![A / 8], m![1], m![A % 8]> = CollectTensor::new(&mut ctx.main, Tensor::zero());
let _o = add_constant(i);
}

f32 Pipeline

vector_narrow_trim() is the Narrow step that converts the tensor from 8-way to 4-way before the float operation. vector_widen_pad() is the Widen step that converts back to 8-way afterward.

#![allow(unused)]
fn main() {
#![feature(adt_const_params)]
extern crate furiosa_opt_std;
use furiosa_opt_std::prelude::*;
axes![A = 512, B = 2];

fn sigmoid<'l, const T: Tu>(
    input: CollectTensor<'l, T, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]>,
) -> VectorFinalTensor<'l, T, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]> {
    input
        .vector_init()
        .vector_intra_slice_tag(TagMode::Zero)
        .vector_narrow_trim::<m![A % 2 # 4]>() // Narrow: Way8 -> Way4
        .vector_fp_unary(FpUnaryOp::Sigmoid)
        .vector_widen_pad::<m![A % 2 # 8]>() // Widen: Way4 -> Way8
        .vector_final()
}

let mut ctx = Context::acquire();

let i: CollectTensor<'_, _, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]> = CollectTensor::new(&mut ctx.main, Tensor::zero());
let _o = sigmoid(i);
}

Bitwise Abs on f32

Clearing the sign bit is \(|x|\), reached with no float ALU: the mask is an ordinary i32 literal because the stream is read as i32 for the one op that needs it. Both reinterprets are free, so this costs exactly one LogicAnd.

#![allow(unused)]
fn main() {
#![feature(adt_const_params)]
extern crate furiosa_opt_std;
use furiosa_opt_std::prelude::*;
axes![A = 512, B = 2];

fn abs<'l, const T: Tu>(
    input: CollectTensor<'l, T, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]>,
) -> VectorFinalTensor<'l, T, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]> {
    input
        .vector_init()
        .vector_intra_slice_tag(TagMode::Zero)
        .vector_reinterpret::<i32>() // read the bits as i32
        .vector_logic(LogicBinaryOpI32::BitAnd, 0x7fff_ffff) // clear the sign bit
        .vector_reinterpret::<f32>() // and back
        .vector_final()
}

let mut ctx = Context::acquire();

let i: CollectTensor<'_, _, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]> = CollectTensor::new(&mut ctx.main, Tensor::zero());
let _o = abs(i);
}

Single-Stream Argument Mode

BinaryArgMode::Mode10 swaps the stream and operand positions, so SubFxp computes operand - stream (here, 7 - x) rather than the default stream - operand.

#![allow(unused)]
fn main() {
#![feature(adt_const_params)]
extern crate furiosa_opt_std;
use furiosa_opt_std::prelude::*;
axes![A = 2048, B = 2];

fn bias_minus_x<'l, const T: Tu>(
    input: CollectTensor<'l, T, i32, m![1], m![B], m![A / 8], m![1], m![A % 8]>,
) -> VectorFinalTensor<'l, T, i32, m![1], m![B], m![A / 8], m![1], m![A % 8]> {
    input
        .vector_init()
        .vector_intra_slice_tag(TagMode::Zero)
        .vector_fxp_with_mode(FxpBinaryOp::SubFxp, BinaryArgMode::Mode10, 7) // compute 7 - x
        .vector_final()
}

let mut ctx = Context::acquire();

let i: CollectTensor<'_, _, i32, m![1], m![B], m![A / 8], m![1], m![A % 8]> = CollectTensor::new(&mut ctx.main, Tensor::zero());
let _o = bias_minus_x(i);
}

VRF Operand

Pre-loaded VRF data as an operand:

#![allow(unused)]
fn main() {
#![feature(adt_const_params)]
extern crate furiosa_opt_std;
use furiosa_opt_std::prelude::*;
axes![A = 2048, B = 2, N = 256];

fn vrf_add<'l, const T: Tu>(
    input: CollectTensor<'l, T, i32, m![1], m![B], m![A / 8], m![N], m![A % 8]>,
    vrf: &VrfTensor<i32, m![1], m![B], m![A / 8], m![A % 8]>,
) -> VectorFinalTensor<'l, T, i32, m![1], m![B], m![A / 8], m![N], m![A % 8]> {
    input
        .vector_init()
        .vector_intra_slice_tag(TagMode::Zero)
        .vector_fxp(FxpBinaryOp::AddFxp, vrf)
        .vector_final()
}

let mut ctx = Context::acquire();

let i: CollectTensor<'_, _, i32, m![1], m![B], m![A / 8], m![N], m![A % 8]> = CollectTensor::new(&mut ctx.main, Tensor::zero());
let v: VrfTensor<i32, m![1], m![B], m![A / 8], m![A % 8]> = VrfTensor::zero();
let _o = vrf_add(i, &v);
}

Stash on the Fp-Only Path

Stash at an early stage, then use it later in a Clip operation. This implements max(2 * x, x):

#![allow(unused)]
fn main() {
#![feature(adt_const_params)]
extern crate furiosa_opt_std;
use furiosa_opt_std::prelude::*;
axes![A = 512, B = 2];

fn residual_max<'l, const T: Tu>(
    input: CollectTensor<'l, T, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]>,
) -> VectorFinalTensor<'l, T, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]> {
    input
        .vector_init()                                        // enter VE
        .vector_intra_slice_tag(TagMode::Zero) // start the intra-slice path
        .vector_stash()                                       // save original x
        .vector_narrow_trim::<m![A % 2 # 4]>()                  // narrow to Way4
        .vector_fp_binary(FpBinaryOp::MulF(FpMulAlu::Mul0), 2.0f32) // compute 2 * x
        .vector_widen_pad::<m![A % 2 # 8]>()                   // widen back to Way8
        .vector_clip(ClipBinaryOpF32::Max, Stash)             // max(2 * x, x)
        .vector_final()
}

let mut ctx = Context::acquire();

let i: CollectTensor<'_, _, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]> = CollectTensor::new(&mut ctx.main, Tensor::zero());
let _o = residual_max(i);
}

Stash on the Fxp-Only Path

Stash at an early stage, then use it later in a Clip operation. This implements max(x + bias, x):

#![allow(unused)]
fn main() {
#![feature(adt_const_params)]
extern crate furiosa_opt_std;
use furiosa_opt_std::prelude::*;
axes![A = 2048, B = 2];

fn stash_at_fxp<'l, const T: Tu>(
    input: CollectTensor<'l, T, i32, m![1], m![B], m![A / 8], m![1], m![A % 8]>,
) -> VectorFinalTensor<'l, T, i32, m![1], m![B], m![A / 8], m![1], m![A % 8]> {
    input
        .vector_init()                                      // enter VE
        .vector_intra_slice_tag(TagMode::Zero) // start the intra-slice path
        .vector_stash()                                     // save original x
        .vector_fxp(FxpBinaryOp::AddFxp, 100)               // compute x + bias
        .vector_clip(ClipBinaryOpI32::Max, Stash)           // compute max(x + bias, x)
        .vector_final()
}

let mut ctx = Context::acquire();

let i: CollectTensor<'_, _, i32, m![1], m![B], m![A / 8], m![1], m![A % 8]> = CollectTensor::new(&mut ctx.main, Tensor::zero());
let _o = stash_at_fxp(i);
}

Stash Across Narrow and Widen

Stash before narrowing, consume after widening. This computes max(sigmoid(x), x):

#![allow(unused)]
fn main() {
#![feature(adt_const_params)]
extern crate furiosa_opt_std;
use furiosa_opt_std::prelude::*;
axes![A = 512, B = 2];

fn stash_across_narrow_widen<'l, const T: Tu>(
    input: CollectTensor<'l, T, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]>,
) -> VectorFinalTensor<'l, T, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]> {
    input
        .vector_init()                                     // enter VE
        .vector_intra_slice_tag(TagMode::Zero) // start the intra-slice path
        .vector_stash()                                    // save x (Way8)
        .vector_narrow_trim::<m![A % 2 # 4]>()               // narrow to Way4
        .vector_fp_unary(FpUnaryOp::Sigmoid)               // compute sigmoid(x) in Way4
        .vector_widen_pad::<m![A % 2 # 8]>()                // widen back to Way8
        .vector_clip(ClipBinaryOpF32::Max, Stash)          // compute max(sigmoid(x), x)
        .vector_final()
}

let mut ctx = Context::acquire();

let i: CollectTensor<'_, _, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]> = CollectTensor::new(&mut ctx.main, Tensor::zero());
let _o = stash_across_narrow_widen(i);
}

Pair Add

Zip two interleaved groups with integer add:

#![allow(unused)]
fn main() {
#![feature(adt_const_params)]
extern crate furiosa_opt_std;
use furiosa_opt_std::prelude::*;
axes![A = 2048, B = 2, I = 2];

fn pair_add<'l, const T: Tu>(
    input: CollectTensor<'l, T, i32, m![1], m![B], m![A / 8], m![I], m![A % 8]>,
) -> VectorFinalTensor<'l, T, i32, m![1], m![B], m![A / 8], m![1], m![A % 8]> {
    input
        .vector_init()
        .vector_intra_slice_unzip::<I, m![1 # 2], m![1]>()
        .vector_clip_zip(ClipBinaryOpI32::AddFxp)
        .vector_final()
}

let mut ctx = Context::acquire();

let i: CollectTensor<'_, _, i32, m![1], m![B], m![A / 8], m![I], m![A % 8]> = CollectTensor::new(&mut ctx.main, Tensor::zero());
let _o = pair_add(i);
}

Pair Per-Side Preprocessing

Asymmetric preprocessing scales only group 0 before zip:

#![allow(unused)]
fn main() {
#![feature(adt_const_params)]
extern crate furiosa_opt_std;
use furiosa_opt_std::prelude::*;
axes![A = 2048, B = 2, I = 2];

fn pair_preprocess_one_side<'l, const T: Tu>(
    input: CollectTensor<'l, T, i32, m![1], m![B], m![A / 8], m![I], m![A % 8]>,
) -> VectorFinalTensor<'l, T, i32, m![1], m![B], m![A / 8], m![1], m![A % 8]> {
    input
        .vector_init()
        .vector_intra_slice_unzip::<I, m![1 # 2], m![1]>()
        .vector_fxp(FxpBinaryOp::MulInt, 10, ())   // group 0 only
        .vector_clip_zip(ClipBinaryOpI32::AddFxp)
        .vector_final()
}

let mut ctx = Context::acquire();

let i: CollectTensor<'_, _, i32, m![1], m![B], m![A / 8], m![I], m![A % 8]> = CollectTensor::new(&mut ctx.main, Tensor::zero());
let _o = pair_preprocess_one_side(i);
}

Pair Float Pipeline with Zip

Both groups traverse the float path (narrow -> fp -> zip -> widen):

#![allow(unused)]
fn main() {
#![feature(adt_const_params)]
extern crate furiosa_opt_std;
use furiosa_opt_std::prelude::*;
axes![A = 512, B = 2, I = 2];

fn pair_fp_mul_zip<'l, const T: Tu>(
    input: CollectTensor<'l, T, f32, m![1], m![B], m![A / 2], m![I], m![A % 2 # 8]>,
) -> VectorFinalTensor<'l, T, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]> {
    input
        .vector_init()
        .vector_intra_slice_unzip::<I, m![1 # 2], m![1]>()
        .vector_narrow_split::<m![1 # 2], m![A % 2 # 4]>()        // both groups: Way8 -> Way4
        .vector_fp_zip(FpBinaryOp::MulF(FpMulAlu::Mul0))   // group0 * group1 (Way4)
        .vector_widen_concat::<m![1], m![A % 2 # 8]>()           // Way4 -> Way8
        .vector_final()
}

let mut ctx = Context::acquire();

let i: CollectTensor<'_, _, f32, m![1], m![B], m![A / 2], m![I], m![A % 2 # 8]> = CollectTensor::new(&mut ctx.main, Tensor::zero());
let _o = pair_fp_mul_zip(i);
}

Pair Per-Group Preprocessing

Apply a different operation to each group before zipping:

#![allow(unused)]
fn main() {
#![feature(adt_const_params)]
extern crate furiosa_opt_std;
use furiosa_opt_std::prelude::*;
axes![A = 512, B = 2, I = 2];

fn pair_asymmetric_preprocess<'l, const T: Tu>(
    input: CollectTensor<'l, T, f32, m![1], m![B], m![A / 2], m![I], m![A % 2 # 8]>,
) -> VectorFinalTensor<'l, T, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]> {
    input
        .vector_init()
        .vector_intra_slice_unzip::<I, m![1 # 2], m![1]>()
        .vector_narrow_split::<m![1 # 2], m![A % 2 # 4]>()
        .vector_fp_unary(FpUnaryOp::Exp, true, false)         // group 0: exp(x), group 1: skip
        .vector_fp_zip(FpBinaryOp::MulF(FpMulAlu::Mul0))   // exp(group0) * group1
        .vector_widen_concat::<m![1], m![A % 2 # 8]>()
        .vector_final()
}

let mut ctx = Context::acquire();

let i: CollectTensor<'_, _, f32, m![1], m![B], m![A / 2], m![I], m![A % 2 # 8]> = CollectTensor::new(&mut ctx.main, Tensor::zero());
let _o = pair_asymmetric_preprocess(i);
}

Pair Zip Argument Mode

BinaryArgMode::Mode10 swaps the two grouped streams when zipping:

#![allow(unused)]
fn main() {
#![feature(adt_const_params)]
extern crate furiosa_opt_std;
use furiosa_opt_std::prelude::*;
axes![A = 512, B = 2, I = 2];

fn pair_sub_reverse<'l, const T: Tu>(
    input: CollectTensor<'l, T, f32, m![1], m![B], m![A / 2], m![I], m![A % 2 # 8]>,
) -> VectorFinalTensor<'l, T, f32, m![1], m![B], m![A / 2], m![1], m![A % 2 # 8]> {
    input
        .vector_init()
        .vector_intra_slice_unzip::<I, m![1 # 2], m![1]>()
        .vector_narrow_split::<m![1 # 2], m![A % 2 # 4]>()
        .vector_fp_zip_with_mode(FpBinaryOp::SubF, BinaryArgMode::Mode10) // compute group1 - group0
        .vector_widen_concat::<m![1], m![A % 2 # 8]>()
        .vector_final()
}

let mut ctx = Context::acquire();

let i: CollectTensor<'_, _, f32, m![1], m![B], m![A / 2], m![I], m![A % 2 # 8]> = CollectTensor::new(&mut ctx.main, Tensor::zero());
let _o = pair_sub_reverse(i);
}

Performance

Throughput is full 8-way (8 elements per cycle) on the Logic, Fxp, and Clip clusters. The Float cluster runs at 4-way, and the Narrow/Widen wrapping around the float path halves effective throughput in practice.

Latency adds one cycle per ALU used. Operations spanning multiple ALUs accumulate their latencies. For example, exp(sqrt(x)) adds 2 cycles (FpFpu for sqrt plus FpExp for exp).