Speed up permute propagation cleanup - #22161
Conversation
Summary: Speed up permute propagation cleanup. ARM lowering spends significant time in permute propagation. Reduce avoidable repeated work while keeping same optimizations: - Canonicalize view/permute chain collection now uses deque + membership set instead of list pop(0) and remove(), avoiding quadratic bookkeeping. - FuseDuplicateUsersPass deduplicates pending producer revisits while preserving same revisit behavior after fusions. - PropagateViewCopyPermutePass still retraces after each moved transform for metadata safety, but defers horizontal/vertical cleanup until full scan finds no more direct propagation moves. Preserves fixed-point behavior while avoiding expensive cleanup after every single moved transform. Behavior-preserving speedup, all existing pass tests pass. Differential Revision: D114224762
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/22161
Note: Links to docs will display an error until the docs builds have been completed. ❌ 8 New Failures, 1 Unclassified FailureAs of commit 858c82b with merge base 469debd ( NEW FAILURES - The following jobs have failed:
UNCLASSIFIED FAILURE - DrCI could not classify the following job because the workflow did not run on the merge base. The failure may be pre-existing on trunk or introduced by this PR:
This comment was automatically generated by Dr. CI and updates every 15 minutes. |
|
@apullin has exported this pull request. If you are a Meta employee, you can view the originating Diff in D114224762. |
This PR needs a
|
Summary:
Speed up permute propagation cleanup.
ARM lowering spends significant time in permute propagation. Reduce avoidable repeated work while keeping same optimizations:
Behavior-preserving speedup, all existing pass tests pass.
Differential Revision: D114224762
cc @digantdesai @freddan80 @per @zingo @oscarandersson8218 @mansnils @Sebastian-Larsson @robell @rascani