Frontend mp flag #384

joerunde · 2024-07-30T20:06:56Z

This adds the --disable-frontend-multiprocessing flag and should also correctly pick up embeddings models to disable the multiprocessing here. (Also some unrelated formatting changes)

The backend stuff is wrapped up in a context manager that handles the process startup and shutdown at exit as well, so that we don't have to muck around much in the existing server lifecycle code

Signed-off-by: Joe Runde <Joseph.Runde@ibm.com>

github-actions · 2024-07-30T20:07:07Z

👋 Hi! Thank you for contributing to the vLLM project.
Just a reminder: PRs would not trigger full CI run by default. Instead, it would only run fastcheck CI which consists a small and essential subset of CI tests to quickly catch errors. You can run other CI tests on top of default ones by unblocking the steps in your fast-check build on Buildkite UI.

Once the PR is approved and ready to go, please make sure to run full CI as it is required to merge (or just use auto-merge).

To run full CI, you can do one of these:

Comment /ready on the PR
Add ready label to the PR
Enable auto-merge.

🚀

robertgshaw2-neuralmagic · 2024-07-30T20:42:57Z

Thanks!

joerunde added 3 commits July 30, 2024 11:13

🚧 wip

5183220

Signed-off-by: Joe Runde <Joseph.Runde@ibm.com>

✨ add flag to disable frontend mp

7ab6bae

Signed-off-by: Joe Runde <Joseph.Runde@ibm.com>

🥅 properly handle shutdown

ff98196

Signed-off-by: Joe Runde <Joseph.Runde@ibm.com>

robertgshaw2-neuralmagic merged commit 453939b into neuralmagic:isolate-oai-server-process Jul 30, 2024
2 checks passed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Frontend mp flag #384

Frontend mp flag #384

joerunde commented Jul 30, 2024

github-actions bot commented Jul 30, 2024

robertgshaw2-neuralmagic commented Jul 30, 2024

Frontend mp flag #384

Frontend mp flag #384

Conversation

joerunde commented Jul 30, 2024

github-actions bot commented Jul 30, 2024

robertgshaw2-neuralmagic commented Jul 30, 2024