)]}'
{
  "commit": "a49335cceab8afb6603152fcc3f7d3b6677366ca",
  "tree": "83f8c06d781a6de77f0b34ec14577bba1e410ac6",
  "parents": [
    "f68f447e8389de9a62e3e80c3c5823cce484c2e5"
  ],
  "author": {
    "name": "Paul Jackson",
    "email": "pj@sgi.com",
    "time": "Tue Sep 06 15:18:09 2005 -0700"
  },
  "committer": {
    "name": "Linus Torvalds",
    "email": "torvalds@g5.osdl.org",
    "time": "Wed Sep 07 16:57:39 2005 -0700"
  },
  "message": "[PATCH] cpusets: oom_kill tweaks\n\nThis patch series extends the use of the cpuset attribute \u0027mem_exclusive\u0027\nto support cpuset configurations that:\n 1) allow GFP_KERNEL allocations to come from a potentially larger\n    set of memory nodes than GFP_USER allocations, and\n 2) can constrain the oom killer to tasks running in cpusets in\n    a specified subtree of the cpuset hierarchy.\n\nHere\u0027s an example usage scenario.  For a few hours or more, a large NUMA\nsystem at a University is to be divided in two halves, with a bunch of student\njobs running in half the system under some form of batch manager, and with a\nbig research project running in the other half.  Each of the student jobs is\nplaced in a small cpuset, but should share the classic Unix time share\nfacilities, such as buffered pages of files in /bin and /usr/lib.  The big\nresearch project wants no interference whatsoever from the student jobs, and\nhas highly tuned, unusual memory and i/o patterns that intend to make full use\nof all the main memory on the nodes available to it.\n\nIn this example, we have two big sibling cpusets, one of which is further\ndivided into a more dynamic set of child cpusets.\n\nWe want kernel memory allocations constrained by the two big cpusets, and user\nallocations constrained by the smaller child cpusets where present.  And we\nrequire that the oom killer not operate across the two halves of this system,\nor else the first time a student job runs amuck, the big research project will\nlikely be first inline to get shot.\n\nTweaking /proc/\u003cpid\u003e/oom_adj is not ideal -- if the big research project\nreally does run amuck allocating memory, it should be shot, not some other\ntask outside the research projects mem_exclusive cpuset.\n\nI propose to extend the use of the \u0027mem_exclusive\u0027 flag of cpusets to manage\nsuch scenarios.  Let memory allocations for user space (GFP_USER) be\nconstrained by a tasks current cpuset, but memory allocations for kernel space\n(GFP_KERNEL) by constrained by the nearest mem_exclusive ancestor of the\ncurrent cpuset, even though kernel space allocations will still _prefer_ to\nremain within the current tasks cpuset, if memory is easily available.\n\nLet the oom killer be constrained to consider only tasks that are in\noverlapping mem_exclusive cpusets (it won\u0027t help much to kill a task that\nnormally cannot allocate memory on any of the same nodes as the ones on which\nthe current task can allocate.)\n\nThe current constraints imposed on setting mem_exclusive are unchanged.  A\ncpuset may only be mem_exclusive if its parent is also mem_exclusive, and a\nmem_exclusive cpuset may not overlap any of its siblings memory nodes.\n\nThis patch was presented on linux-mm in early July 2005, though did not\ngenerate much feedback at that time.  It has been built for a variety of\narch\u0027s using cross tools, and built, booted and tested for function on SN2\n(ia64).\n\nThere are 4 patches in this set:\n  1) Some minor cleanup, and some improvements to the code layout\n     of one routine to make subsequent patches cleaner.\n  2) Add another GFP flag - __GFP_HARDWALL.  It marks memory\n     requests for USER space, which are tightly confined by the\n     current tasks cpuset.\n  3) Now memory requests (such as KERNEL) that not marked HARDWALL can\n     if short on memory, look in the potentially larger pool of memory\n     defined by the nearest mem_exclusive ancestor cpuset of the current\n     tasks cpuset.\n  4) Finally, modify the oom killer to skip any task whose mem_exclusive\n     cpuset doesn\u0027t overlap ours.\n\nPatch (1), the one time I looked on an SN2 (ia64) build, actually saved 32\nbytes of kernel text space.  Patch (2) has no affect on the size of kernel\ntext space (it just adds a preprocessor flag).  Patches (3) and (4) added\nabout 600 bytes each of kernel text space, mostly in kernel/cpuset.c, which\nmatters only if CONFIG_CPUSET is enabled.\n\nThis patch:\n\nThis patch applies a few comment and code cleanups to mm/oom_kill.c prior to\napplying a few small patches to improve cpuset management of memory placement.\n\nThe comment changed in oom_kill.c was seriously misleading.  The code layout\nchange in select_bad_process() makes room for adding another condition on\nwhich a process can be spared the oom killer (see the subsequent\ncpuset_nodes_overlap patch for this addition).\n\nAlso a couple typos and spellos that bugged me, while I was here.\n\nThis patch should have no material affect.\n\nSigned-off-by: Paul Jackson \u003cpj@sgi.com\u003e\nSigned-off-by: Andrew Morton \u003cakpm@osdl.org\u003e\nSigned-off-by: Linus Torvalds \u003ctorvalds@osdl.org\u003e\n",
  "tree_diff": [
    {
      "type": "modify",
      "old_id": "1e56076672f5870e9766753270d77442e5743ac9",
      "old_mode": 33188,
      "old_path": "mm/oom_kill.c",
      "new_id": "3a1d46502938eaf7b5c03216a3f21a53e4d54deb",
      "new_mode": 33188,
      "new_path": "mm/oom_kill.c"
    }
  ]
}
