You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: experimental/task_scheduling/README.md
+76-1Lines changed: 76 additions & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -30,7 +30,82 @@ The manager also computes a warp-group-rounded initial register budget and threa
30
30
31
31
Automatic value routing hides storage positions from callback authors. Each callback receives its routed inputs after the `StageInfo` argument and returns the values declared by its captured work method. The frozen `DeviceStep` performs the corresponding value-stack operations, while `DeviceTask.make_context(tasks_inputs)` creates the immutable execution context. The tuple representation remains an internal compiler detail; no placeholder or named routing slot is exposed to resource authors.
32
32
33
-
`outputs=N` creates `N` independent lexical `ScheduleValue` instances. A work call returns one value directly or a tuple that can be unpacked normally, and its device callback returns the same number of runtime values. A value produced inside a conditional or domain loop cannot escape that scope accidentally. A domain loop(don't support while senmantic and contional early break) that needs state across iterations uses `domain_loop(start, end, step, body, *initial_values)`, with one callback parameter per initial value. The schedule builder invokes `body(*iter_values)` once while capturing the loop body, so every resource call remains a schedule step. The callback must return one work-call output per input; those outputs become both the next-iteration values and the post-loop results. The device adapter preserves zero-trip pass-through and materializes the backedge without exposing routing positions to callbacks. Work methods read the current device offset from `stage_info.loop_offset`; it is not a callback parameter or a routed value. Loop bodies may use `first_iter()`, `last_iter()`, and `every(period, start=...)` to capture iteration-guarded regions.
33
+
`outputs=N` creates `N` independent lexical `ScheduleValue` instances. A work call returns one value directly or a tuple that can be unpacked normally, and its device callback returns the same number of runtime values. A value produced inside a conditional or domain loop cannot escape that scope accidentally. A domain loop that needs state across iterations uses `domain_loop(start, end, step, body, *initial_values)`, with one callback parameter per initial value. The schedule builder invokes `body(*iter_values)` once while capturing the loop body, so every resource call remains a schedule step. The callback must return one work-call output per input; those outputs become both the next-iteration values and the post-loop results. The device adapter preserves zero-trip pass-through and materializes the backedge without exposing routing positions to callbacks. Work methods read the current device offset from `stage_info.loop_offset`; it is not a callback parameter or a routed value. Loop bodies may use `first_iter()`, `last_iter()`, and `every(period, start=...)` to capture iteration-guarded regions.
34
+
35
+
`when_true()` and `when_false()` regions may nest, including inside iteration
36
+
guards. An inner region executes only when every enclosing guard is active.
37
+
Predicates and other work outputs produced inside a branch may be used by its
38
+
nested regions, but cannot escape to a parent or sibling scope. Reusing the same
39
+
routed predicate preserves true/false correlation in host analysis.
40
+
41
+
The runnable `tutorial/01_copy_basics/04_copy_tma_nested_conditional.py` example
42
+
extends the conditional TMA copy with all four two-level true/false combinations
43
+
and a third-level condition. It checks the copied tensor and exact per-row branch
44
+
markers, including the absence of writes from inactive branches:
`break_loop()` exits the enclosing domain loop immediately, including from a
51
+
nested conditional. It skips the rest of that iteration and resumes after the
52
+
loop. The context-manager handle provides the same operation:
53
+
54
+
```python
55
+
with ts.domain_loop(num_tiles) as loop:
56
+
done = resource.is_done()
57
+
with ts.when_true(done):
58
+
loop.break_loop()
59
+
resource.process()
60
+
```
61
+
62
+
Use `ts.break_loop()` in a functional loop body. A bare break skips the normal
63
+
backedge: returned loop-carried results retain their values from the last
64
+
completed iteration (or the initial values if the first iteration breaks).
65
+
Pass explicit results to return values computed during the interrupted iteration:
66
+
67
+
```python
68
+
initial = resource.init_state()
69
+
70
+
defbody(state):
71
+
updated = resource.advance_state(state)
72
+
done = resource.is_done(updated)
73
+
with ts.when_true(done):
74
+
ts.break_loop(updated)
75
+
return updated
76
+
77
+
result = ts.domain_loop(0, num_tiles, 1, body, initial)
78
+
resource.consume_state(result)
79
+
```
80
+
81
+
For multiple carried inputs, use `ts.break_loop(next_count, next_total)` in the
82
+
same order as the functional loop's inputs. Supply all carried results or none.
83
+
Each explicit argument must be a visible routed `ScheduleValue` with the same
84
+
pipeline-stage provenance and a compatible CUDA type as its carried input;
85
+
ordinary constants must first be returned by a work method. Values created
86
+
inside a nested branch may be returned directly by a break in that branch.
87
+
The exit values are captured before branch-local routes are discarded. Wrong
88
+
arity, foreign-schedule values, and values escaping a sibling/child scope are
89
+
rejected during capture; CUDA types are checked during device compilation.
90
+
An untaken break uses the normal body return values, and a zero-trip loop still
91
+
returns its initial values.
92
+
93
+
Side effects and pipeline-state updates performed before the break are retained.
94
+
A break exits only the domain loop, so an enclosing `work_tile_loop()` continues.
95
+
`last_iter()` still means the last iteration of the declared range; it is not an
96
+
exit hook. Breaks outside a domain loop and expired loop handles are rejected.
97
+
Unbounded while loops remain unsupported.
98
+
99
+
The nested-copy tutorial also accepts `--stop-row 5`. Both producer and consumer
100
+
break before acquiring/waiting for row 5, and the example verifies that the
101
+
uncopied output and trace remain zero. A pipeline schedule must ensure that its
102
+
producer and consumer exit consistently and finish any acquired work before
103
+
exiting; `break_loop()` does not implicitly release or drain an interrupted step.
104
+
105
+
Host expansion honors breaks for its selected opaque-condition assignment and
106
+
representative loop bounds, including zero-trip loops. Opaque assignments remain
107
+
fixed during each explored execution; these checks do not prove safety for all
108
+
possible iteration-varying runtime predicate sequences.
34
109
35
110
`StageInfo` contains the current `stage_idx`, `phase`, selected full `barrier`, zero-based iteration count, loop offset/bounds, work label, and owning `ExecutionContext`; loop offset/bounds are `None` for peeled work outside a domain loop. This lets pipeline payload work use `stage_info.barrier` without knowing how barrier arrays are stored by the device manager.
0 commit comments