[Feature][habana-main]: HPUAttentionImpl support for 'Encoder_Decoder' #370

xuechendi · 2024-10-07T19:45:21Z

🚀 The feature, motivation and pitch

Llama3.2 vision (Mllama) models requires model runner as "Enocoder_Decoder_Model_Runner"
which includes:

prepare "encoder_seq_lens" and "encoder_seq_lens_tensor" when preparing input data
necessary fix for "HPUModelRunner - prepare_input_tensors" xuechendi@1f5a702
enable "Encoder self-attention" and "encoder/decoder cross-attention" in HPUAttentionImpl

test cmd:

python offline_inference_vision_language.py --model_type mllama

Error msg:

File "/workspace/vllm/vllm/attention/backends/hpu_attn.py", line 159, in forward
[rank0]:     raise NotImplementedError("Encoder self-attention and "
[rank0]: NotImplementedError: Encoder self-attention and encoder/decoder cross-attention are not implemented for HPUAttentionImpl

Alternatives

No response

Additional context

No response

Before submitting a new issue...

Make sure you already searched for relevant issues, and asked the chatbot living at the bottom right corner of the documentation page, which can answer lots of frequently asked questions.

The text was updated successfully, but these errors were encountered:

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[Feature][habana-main]: HPUAttentionImpl support for 'Encoder_Decoder' #370

[Feature][habana-main]: HPUAttentionImpl support for 'Encoder_Decoder' #370

xuechendi commented Oct 7, 2024 •

edited

Loading

[Feature][habana-main]: HPUAttentionImpl support for 'Encoder_Decoder' #370

[Feature][habana-main]: HPUAttentionImpl support for 'Encoder_Decoder' #370

Comments

xuechendi commented Oct 7, 2024 • edited Loading

🚀 The feature, motivation and pitch

Alternatives

Additional context

Before submitting a new issue...

xuechendi commented Oct 7, 2024 •

edited

Loading