forked from torvalds/linux
-
Notifications
You must be signed in to change notification settings - Fork 7
Commit
This commit does not belong to any branch on this repository, and may belong to a fork outside of the repository.
kernel/cgroup: Add "dmem" memory accounting cgroup
This code is based on the RDMA and misc cgroup initially, but now uses page_counter. It uses the same min/low/max semantics as the memory cgroup as a result. There's a small mismatch as TTM uses u64, and page_counter long pages. In practice it's not a problem. 32-bits systems don't really come with >=4GB cards and as long as we're consistently wrong with units, it's fine. The device page size may not be in the same units as kernel page size, and each region might also have a different page size (VRAM vs GART for example). The interface is simple: - Call dmem_cgroup_register_region() - Use dmem_cgroup_try_charge to check if you can allocate a chunk of memory, use dmem_cgroup__uncharge when freeing it. This may return an error code, or -EAGAIN when the cgroup limit is reached. In that case a reference to the limiting pool is returned. - The limiting cs can be used as compare function for dmem_cgroup_state_evict_valuable. - After having evicted enough, drop reference to limiting cs with dmem_cgroup_pool_state_put. This API allows you to limit device resources with cgroups. You can see the supported cards in /sys/fs/cgroup/dmem.capacity You need to echo +dmem to cgroup.subtree_control, and then you can partition device memory. Co-developed-by: Friedrich Vock <friedrich.vock@gmx.de> Signed-off-by: Friedrich Vock <friedrich.vock@gmx.de> Co-developed-by: Maxime Ripard <mripard@kernel.org> Signed-off-by: Maarten Lankhorst <dev@lankhorst.se> Acked-by: Tejun Heo <tj@kernel.org> Link: https://lore.kernel.org/r/20241204143112.1250983-1-dev@lankhorst.se Signed-off-by: Maxime Ripard <mripard@kernel.org>
- Loading branch information
Showing
11 changed files
with
1,060 additions
and
10 deletions.
There are no files selected for viewing
This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Original file line number | Diff line number | Diff line change |
---|---|---|
@@ -0,0 +1,9 @@ | ||
================== | ||
Cgroup Kernel APIs | ||
================== | ||
|
||
Device Memory Cgroup API (dmemcg) | ||
========================= | ||
.. kernel-doc:: kernel/cgroup/dmem.c | ||
:export: | ||
|
This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Original file line number | Diff line number | Diff line change |
---|---|---|
@@ -0,0 +1,54 @@ | ||
================================== | ||
Long running workloads and compute | ||
================================== | ||
|
||
Long running workloads (compute) are workloads that will not complete in 10 | ||
seconds. (The time let the user wait before he reaches for the power button). | ||
This means that other techniques need to be used to manage those workloads, | ||
that cannot use fences. | ||
|
||
Some hardware may schedule compute jobs, and have no way to pre-empt them, or | ||
have their memory swapped out from them. Or they simply want their workload | ||
not to be preempted or swapped out at all. | ||
|
||
This means that it differs from what is described in driver-api/dma-buf.rst. | ||
|
||
As with normal compute jobs, dma-fence may not be used at all. In this case, | ||
not even to force preemption. The driver with is simply forced to unmap a BO | ||
from the long compute job's address space on unbind immediately, not even | ||
waiting for the workload to complete. Effectively this terminates the workload | ||
when there is no hardware support to recover. | ||
|
||
Since this is undesirable, there need to be mitigations to prevent a workload | ||
from being terminated. There are several possible approach, all with their | ||
advantages and drawbacks. | ||
|
||
The first approach you will likely try is to pin all buffers used by compute. | ||
This guarantees that the job will run uninterrupted, but also allows a very | ||
denial of service attack by pinning as much memory as possible, hogging the | ||
all GPU memory, and possibly a huge chunk of CPU memory. | ||
|
||
A second approach that will work slightly better on its own is adding an option | ||
not to evict when creating a new job (any kind). If all of userspace opts in | ||
to this flag, it would prevent cooperating userspace from forced terminating | ||
older compute jobs to start a new one. | ||
|
||
If job preemption and recoverable pagefaults are not available, those are the | ||
only approaches possible. So even with those, you want a separate way of | ||
controlling resources. The standard kernel way of doing so is cgroups. | ||
|
||
This creates a third option, using cgroups to prevent eviction. Both GPU and | ||
driver-allocated CPU memory would be accounted to the correct cgroup, and | ||
eviction would be made cgroup aware. This allows the GPU to be partitioned | ||
into cgroups, that will allow jobs to run next to each other without | ||
interference. | ||
|
||
The interface to the cgroup would be similar to the current CPU memory | ||
interface, with similar semantics for min/low/high/max, if eviction can | ||
be made cgroup aware. | ||
|
||
What should be noted is that each memory region (tiled memory for example) | ||
should have its own accounting. | ||
|
||
The key is set to the regionid set by the driver, for example "tile0". | ||
For the value of $card, we use drmGetUnique(). |
This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Original file line number | Diff line number | Diff line change |
---|---|---|
@@ -0,0 +1,66 @@ | ||
/* SPDX-License-Identifier: MIT */ | ||
/* | ||
* Copyright © 2023-2024 Intel Corporation | ||
*/ | ||
|
||
#ifndef _CGROUP_DMEM_H | ||
#define _CGROUP_DMEM_H | ||
|
||
#include <linux/types.h> | ||
#include <linux/llist.h> | ||
|
||
struct dmem_cgroup_pool_state; | ||
|
||
/* Opaque definition of a cgroup region, used internally */ | ||
struct dmem_cgroup_region; | ||
|
||
#if IS_ENABLED(CONFIG_CGROUP_DMEM) | ||
struct dmem_cgroup_region *dmem_cgroup_register_region(u64 size, const char *name_fmt, ...) __printf(2,3); | ||
void dmem_cgroup_unregister_region(struct dmem_cgroup_region *region); | ||
int dmem_cgroup_try_charge(struct dmem_cgroup_region *region, u64 size, | ||
struct dmem_cgroup_pool_state **ret_pool, | ||
struct dmem_cgroup_pool_state **ret_limit_pool); | ||
void dmem_cgroup_uncharge(struct dmem_cgroup_pool_state *pool, u64 size); | ||
bool dmem_cgroup_state_evict_valuable(struct dmem_cgroup_pool_state *limit_pool, | ||
struct dmem_cgroup_pool_state *test_pool, | ||
bool ignore_low, bool *ret_hit_low); | ||
|
||
void dmem_cgroup_pool_state_put(struct dmem_cgroup_pool_state *pool); | ||
#else | ||
static inline __printf(2,3) struct dmem_cgroup_region * | ||
dmem_cgroup_register_region(u64 size, const char *name_fmt, ...) | ||
{ | ||
return NULL; | ||
} | ||
|
||
static inline void dmem_cgroup_unregister_region(struct dmem_cgroup_region *region) | ||
{ } | ||
|
||
static inline int dmem_cgroup_try_charge(struct dmem_cgroup_region *region, u64 size, | ||
struct dmem_cgroup_pool_state **ret_pool, | ||
struct dmem_cgroup_pool_state **ret_limit_pool) | ||
{ | ||
*ret_pool = NULL; | ||
|
||
if (ret_limit_pool) | ||
*ret_limit_pool = NULL; | ||
|
||
return 0; | ||
} | ||
|
||
static inline void dmem_cgroup_uncharge(struct dmem_cgroup_pool_state *pool, u64 size) | ||
{ } | ||
|
||
static inline | ||
bool dmem_cgroup_state_evict_valuable(struct dmem_cgroup_pool_state *limit_pool, | ||
struct dmem_cgroup_pool_state *test_pool, | ||
bool ignore_low, bool *ret_hit_low) | ||
{ | ||
return true; | ||
} | ||
|
||
static inline void dmem_cgroup_pool_state_put(struct dmem_cgroup_pool_state *pool) | ||
{ } | ||
|
||
#endif | ||
#endif /* _CGROUP_DMEM_H */ |
This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Oops, something went wrong.