[Model] Jamba support #4115

mzusman · 2024-04-16T12:54:58Z

Add Jamba support to vLLM,
This PR comprises three parts:

The Jamba modeling file that encapsulates the Jamba model weights and logic itself and the mamba cache management.
Passage of the requests ids of the sequence groups and their sequence ids into the modeling file in order to be able to manage the cache.
Passage of the finished request ids into the modeling file as well in order the clean the allocated cache on finished requests

BA-78554: Jurassic 2.5 * worked on jurasic2.5 configuration file, updated jurassic2_5 modeling file to support alternating experts/attn layers * finished working the forward pass of jurassic3.py * finished working the forward pass of jurassic3.py * finished working the forward pass of jurassic3.py * jurassic_3 modeling file works, uses dummy weights initialized by "dummy" flag. Tokenizer raises issues, for now copying the mixtral tokenizer * changed default tokenizer vocab values, loading of custom .pt weight files works. * removed notebook * merging master to jurassic-2.5 to reset head * Merge branch 'master' into jurassic-2.5 * align to master Approved-by: Tomer Asida Approved-by: Mor Zusman

BA-78760: Jamba * Add support for n concat and splitting * change naming * input_metadata is a dict list now in order to pass "n" * clean up code from unecessary changes and prints * Remove kv cache allocation in case of mamba layer * Add the considerations of mamba layer cache into the num of blocks calculation * Delete mamba cache after profile * Remove prints * Cleaning * - and not _ for requirements Approved-by: Tomer Asida

* Remove assertion * adapting jamba vllm to changes after hf release, working on weight loading in modeling file * splitting the JambaDecoderLayer to JambaMambaDecoderLayer and JambaAttentionDecoderLayer * weight loading from hf checkpoint supposedly works, might be a mixup in the MoE between the gated and non-gated weights * Add mamba from jamba modeling file * Remove slow forward * Modifications to mamba_mixer * Save changes, WIP * Fix cache placement * Debugging * Additions and logging * Jamba with mamba cache handling * Clean up * Another cleanup * Use vllm's RMSNorm instead of JambaRMSNorm, Thier implementation is with fused kernel * Clean up and orginization of the objects to handle the mamba cache * Shorten the code for kv cache mem * Move cache handling inside the Mixer * Add mamba to the wheel requirements * Add mamba to the requirements script * Add mamba_metadata * Add to __init__ __all__ * Revert 2 commits ad1a3db 'Add mamba to the requirements script' 75ed2c8 'Add mamba to the wheel requirements' * Clean up * Naming * Apply whitespace suggestions from code review * pass tie_word_embeddings to PretrainedConfig init * Replace repeat with expand as expand doesn't require more mem * Allocate really small cache if needed , don't use meta * Fix for expanded --------- Co-authored-by: Mor Zusman <morz@ai21.com> Co-authored-by: Erez Schwartz <erezs@ai21.com> Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com>

* Drop indecies when finish * min 1 attention layer * CG is working on forward pass passing * Remove comments * cosmetics - rename indecies -> indices, organize some whitespaces * Add some TODOs * Adding mamba cache for cg * Remove useless vars from input_metadata * Remove unused import * Set the seqlen offset to boolean * Return only hidden state * Return only hidden states * Add padding to match forward pass bs * Is prompt instead of seqlen offset * Remove mamba cache class (not used) * Another remove * Remove * Use mamba4gc * Fix mamba forward, run update only on non prompt * Use 1 index after the maximal index * Remove import * Remove import * typo * typo * place holder * Padding and empty token takes it from the first empty place * reformat * Apply suggestions from code review Whitespaces --------- Co-authored-by: Mor Zusman <morz@ai21.com> Co-authored-by: Tomer Asida <tomera@ai21.com> Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com>

Co-authored-by: Mor Zusman <morz@ai21.com>

* Return support for other models apart from jamba * Support n>1 * A little cleanup * Rename * Apply whitespace suggestions from code review * Add max batch size to the main func * Fixed attention kv cache bug * log where requests id are deleted from the dict to debug mode * Fix typo * Align with v0.3.3 vllm code * Remove comments * Take out model config from CUDAGraph object * Fix * Fix typo * Make the kv cache selection cleaner * Another typo * Took the num layers calc outside * Remove the -1 * Set as num layer / period --------- Co-authored-by: Mor Zusman <morz@ai21.com> Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com>

* Return support for other models apart from jamba * Support n>1 * Revert 2 commits d054737 'Support n>1' b5167cc 'Return support for other models apart from jamba' * TP on input and output * Basic TP impl , working, correctness not working * TP is working * Roll back the verification that everything in the weights fits into the model * Cleanup * Use world size func * clean up * Import * Apply whitespace suggestions from code review * Organize imports * Add comment on the unsqueeze in conv1d * Organize and remove redundant code in forward pass * Remove print * Add comments Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com> * White spaces * Set as A * better comment --------- Co-authored-by: Mor Zusman <morz@ai21.com> Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com>

robertgshaw2-redhat · 2024-04-17T01:00:26Z

Cool!

Signed-off-by: Muralidhar Andoorveedu <muralidhar.andoorveedu@centml.ai>

mzusman · 2024-07-02T22:19:51Z

Tests failed due to timeouts to HF
Ready to be merged

Signed-off-by: Muralidhar Andoorveedu <muralidhar.andoorveedu@centml.ai> Co-authored-by: Erez Schwartz <erezs@ai21.com> Co-authored-by: Mor Zusman <morz@ai21.com> Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com> Co-authored-by: Tomer Asida <tomera@ai21.com> Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-authored-by: Muralidhar Andoorveedu <muralidhar.andoorveedu@centml.ai>

mmoskal · 2024-07-09T17:00:33Z

vllm/engine/llm_engine.py

@@ -855,7 +857,7 @@ def step(self) -> List[Union[RequestOutput, EmbeddingRequestOutput]]:
                blocks_to_copy=scheduler_outputs.blocks_to_copy,
                num_lookahead_slots=scheduler_outputs.num_lookahead_slots,
                running_queue_size=scheduler_outputs.running_queue_size,
-            )
+                finished_requests_ids=finished_requests_ids)


@mzusman If scheduler_output.is_empty() it seems the finished request ids would be forgotten (and never actually freed). I guess you may want to call get_and_reset_finished_requests_ids() inside of the if statement?

That's right 👍 Nice catch, I'll open a PR to fix it #6266

I have now realized that this code path is not hit it seems, at least in common circumstances due to this https://github.com/vllm-project/vllm/blob/main/vllm/engine/async_llm_engine.py#L563 but probably good to fix anyways!

Signed-off-by: Muralidhar Andoorveedu <muralidhar.andoorveedu@centml.ai> Co-authored-by: Erez Schwartz <erezs@ai21.com> Co-authored-by: Mor Zusman <morz@ai21.com> Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com> Co-authored-by: Tomer Asida <tomera@ai21.com> Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-authored-by: Muralidhar Andoorveedu <muralidhar.andoorveedu@centml.ai>

Signed-off-by: Muralidhar Andoorveedu <muralidhar.andoorveedu@centml.ai> Co-authored-by: Erez Schwartz <erezs@ai21.com> Co-authored-by: Mor Zusman <morz@ai21.com> Co-authored-by: tomeras91 <57313761+tomeras91@users.noreply.github.com> Co-authored-by: Tomer Asida <tomera@ai21.com> Co-authored-by: Zhuohan Li <zhuohan123@gmail.com> Co-authored-by: Muralidhar Andoorveedu <muralidhar.andoorveedu@centml.ai> Signed-off-by: Alvant <alvasian@yandex.ru>

ErezSC42 and others added 28 commits April 16, 2024 10:13

dtype (vllm-project#6)

00bce1f

Co-authored-by: Mor Zusman <morz@ai21.com>

After merge fixes

30e6dcd

Clean up

5c0efdc

Add release mamba cache to executor_base

19f11f3

Add jamba modifications

1fb817a

Add minimun 1 attention layer

30ae4a1

More fixes

7bd9c0a

Delete mamba cache

d5ac8e8

Jamba padding to the left

60b49b5

Clean up

c583fe8

Add import

c951b7d

Another clean up

da6d0f2

Align to main

eb79923

Fix reduce

919edba

Another fix

4668566

Black format for jamba

11a0737

Formatting

7e3415e

Formatting with format.sh

adbd2ae

Adding to docs and more

6daf2a2

Add to readme

7ee927b

Adding comments for prefill mamba

87fa299

Formating

8bca3b6

mzusman mentioned this pull request Apr 16, 2024

[New Model]: Jamba (MoE Mamba from AI21) #3690

Closed

andoorve and others added 9 commits July 2, 2024 04:52

Address Nick nits and fix CUDAGraph correctness

c92257c

Signed-off-by: Muralidhar Andoorveedu <muralidhar.andoorveedu@centml.ai>

Merge branch 'pipeline-parallel' into jamba-support-pr

60bb1a7

Merge branch 'gh-main' into jamba-support-pr

2ea2b80

Formating and fixing llm engine

10d8f3c

Align with main and format

1331a8f

Fix bug

21c92b4

Format

726ccad

Add intermediate tensors

4b6a491

Format

da5d94a

zhuohan123 enabled auto-merge (squash) July 2, 2024 22:22

zhuohan123 merged commit 9d6a8da into vllm-project:main Jul 2, 2024
70 checks passed

ywang96 mentioned this pull request Jul 3, 2024

[Bugfix] Fix compute_logits in Jamba #6093

Merged

sroy745 mentioned this pull request Jul 8, 2024

[Speculative Decoding] Enabling bonus token in speculative decoding for KV cache based models #5765

Merged

DarkLight1337 mentioned this pull request Jul 9, 2024

I want to add mamba_chat (2.8b) model #2001

Closed

mmoskal reviewed Jul 9, 2024

View reviewed changes

This was referenced Jul 9, 2024

[BugFix] get_and_reset only when scheduler outputs are not empty #6266

Merged

[BugFix][Model] Jamba - Handle aborted requests, Add tests and fix cleanup bug #6425

Merged

youkaichao mentioned this pull request Jul 27, 2024

[core][misc] improve free_finished_seq_groups #6865

Merged

This was referenced Aug 12, 2024

Change interface to causal_conv1d_update for continuous batching Dao-AILab/causal-conv1d#29

Merged

Change interface to selective_state_update for continuous batching state-spaces/mamba#521

Merged

mzusman mentioned this pull request Aug 19, 2024

[Kernel/Model] Migrate mamba_ssm and causal_conv1d kernels to vLLM #7651

Merged

nivibilla mentioned this pull request Aug 23, 2024

[Feature] Jamba 1.5 Support PLS sgl-project/sglang#1190

Open

2 tasks

This was referenced Aug 29, 2024

[Kernel] Change interface to Mamba causal_conv1d_update for continuous batching #8012

Merged

[Kernel] Change interface to Mamba selective_state_update for continuous batching #8039

Merged

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[Model] Jamba support #4115

[Model] Jamba support #4115

mzusman commented Apr 16, 2024 •

edited by DarkLight1337

Loading

robertgshaw2-redhat commented Apr 17, 2024

mzusman commented Jul 2, 2024

mmoskal Jul 9, 2024

mzusman Jul 9, 2024 •

edited

Loading

mmoskal Jul 9, 2024

[Model] Jamba support #4115

[Model] Jamba support #4115

Conversation

mzusman commented Apr 16, 2024 • edited by DarkLight1337 Loading

robertgshaw2-redhat commented Apr 17, 2024

mzusman commented Jul 2, 2024

mmoskal Jul 9, 2024

Choose a reason for hiding this comment

mzusman Jul 9, 2024 • edited Loading

Choose a reason for hiding this comment

mmoskal Jul 9, 2024

Choose a reason for hiding this comment

mzusman commented Apr 16, 2024 •

edited by DarkLight1337

Loading

mzusman Jul 9, 2024 •

edited

Loading