perf: optimize generic tensor conversion - #832
Draft
baominghelly wants to merge 1 commit into
Draft
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
TensorFromPybind11Handle.Motivation
TensorFromPybind11Handleruns for every tensor argument passed through the generated eager Python bindings. The previous implementation repeatedly created attribute-name objects and converted Python strings/sequences through generic pybind11 casters.On an NVIDIA A100-SXM4-80GB, an alternating A/B run of the exact patch used FP32 16x16 GEMM with
implementation_index=0, 200 warmups, 2,000 calls per repetition, 15 repetitions per process, and four runs per variant:Type of Change
feat— new feature / new operator / new platformfix— bug fixperf— performance improvement (no behavioral change)refactor— code restructuring without behavior changetest— adding or fixing tests onlydocs— documentation onlybuild/ci— build system or CI configurationchore— tooling, formatting, or other non-code changes!in the Conventional Commits prefix or aBREAKING CHANGE:footer)Platforms Affected
WITH_CPU)WITH_NVIDIA)WITH_ILUVATAR)WITH_METAX)WITH_CAMBRICON)WITH_MOORE)WITH_ASCEND)WITH_TORCH)Smoke Test Result
The two failures are the existing FP32 GEMM tolerance cases for implementations 0 and 1. Both failures reproduce with the same output values on an independently built, unmodified
masterat26228d8; the candidate and baseline GEMM smoke subsets each report2 failed, 8 passed.Test Results on Supported Platforms
Full `pytest` output (optional)
Benchmark / Performance Impact
Notes for Reviewers
data_ptr,shape,dtype,device, andstride).THPVariable_Unpack, Torch wrapper generation changes, CMake toggles, profiling hooks, benchmark scripts, or profiling logs.