|
|
|
I'm currently a senior software engineer working at AMD ROCm (previously at Huawei Ascend), building vLLM inference engine for GPU/NPU software ecosystem.
I'm also a maintainer (member) of vllm-project, mainly focusing on:
- Performance optimization & bugfix for AMD Instinct GPUs.
- Multi-modality inference (especially ViT execution optimization).
- Hardware portability (especially OOT device extensibility).




