# OpenIFS LBUD23 tendencies --\> out-of-mem

**URL:** <https://forum.ecmwf.int/t/openifs-lbud23-tendencies-out-of-mem/12175>\
**Category:** OpenIFS\
**Tags:** openifs-configuration-and-namelists, openifs-43r3, openifs-model-output\
**Created:** [28 April 2022 08:44 UTC](https://forum.ecmwf.int/t/openifs-lbud23-tendencies-out-of-mem/12175 "2022-04-28T08:44:32Z")\
**Posts on this page:** 6\
**Page:** 1

<div class="post-metadata">

**Author:** ![Thomas\_Batelaan](https://forum.ecmwf.int/letter_avatar_proxy/v4/letter/t/ee7513/32.png) [@Thomas\_Batelaan](https://forum.ecmwf.int/u/Thomas_Batelaan)\
**Post date:** [28 April 2022 08:44 UTC](https://forum.ecmwf.int/t/openifs-lbud23-tendencies-out-of-mem/12175/1 "2022-04-28T08:44:32Z")

</div>

Hi all,

I try to run a T1279 case and output PEXTRA tendencies, but I am overwriting my memory. I can run the model with regular (non-PEXTRA) grib codes, but I am struggling to output the tendencies.

I share hereby the relevant fields from the namelist:

> &NAEPHY  
> LEPHYS=true,&nbsp; &nbsp; &nbsp; ! switch the full ECMWF physics package on/off.  
> LBUD23=true,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! enable computation of physics tendencies and budget diagnostics  
> LEVDIF=true,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp; ! turn on/off the vertical diffusion scheme.  
> LESURF=true,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; ! turn on/off the interface surface processes.  
> LECOND=true,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! turn on/off the large-scale condensation processes.  
> LECUMF=true,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! turn on/off the mass-flux cumulus convection.  
> LEPCLD=true,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; ! turn on/off the prognostic cloud scheme.  
> LEEVAP=true,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;&nbsp; ! turn on/off the evaporation of precipitation  
> LEVGEN=true,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; ! turn on/off Van Genuchten hydrology (with soil type field)  
> LESSRO=true,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! turn on/off orographic (VIC-type) runoff  
> LECURR=false, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! if true, ocean current boundary condition is used.  
> LEOCWA=true,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! WARM OCEAN LAYER PARAMETRIZATION  
> LEGWDG=true,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! turn on/off gravity wave drag.  
> LEGWWMS=true, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! turn on/off the Warner-McIntyre-Scinocca non-orographic gravity wave drag scheme.  
> LEOZOC=false, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! turn on/off the climatological ozone.  
> LEQNGT=true,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! turn on/off the negative humidity fixer.  
> LERADI=true,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! turn on/off the radiation scheme.  
> LERADS=true,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! turn on/off the interactive surface radiative properties.  
> LESICE=true,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! turn on/off the interactive sea-ice processes.  
> LEO3CH=true,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! turn on/off the ozone chemistry (for prognostic ozone).  
> LEDCLD=true,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! turn on/off the diagnostic cloud scheme.  
> LDUCTDIA=false, &nbsp; &nbsp; &nbsp; &nbsp; ! turn on/off computation and archiving of ducting diagnostics.  
> LELIGHT=false,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! ACTIVATES LIGHTNING PARAMETRIZATION  
> LWCOU=true, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! turn on/off coupled wave model (n.b. always off for OpenIFS model version 38r1).  
> LWCOU2W=true, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! turn on/off two-way interaction with the wave model (n.b. always off for OpenIFS model version 38r1).  
> NSTPW=1,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! INTEGER&nbsp; &nbsp; FREQUENCY OF CALL TO THE WAVE MODEL.  
> RDEGREW=0.5,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! RESOLUTION OF THE WAVE MODEL (DEGREES).  
> RSOUTW=-81.0, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! SOUTH BOUNDARY OF THE WAVE MODEL.  
> RNORTW=81.0,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! NORTH BOUNDARY OF THE WAVE MODEL.  
> /  
> &NAMFPC  
> CFPFMT="MODEL",  
> !  
> !&nbsp; output on model levels  
> NFP3DFS=6,   
> MFP3DFS(:)=93,95,98,102,105,109,   
> NRFP3S(:)=1, ! I also tried it with all the 137 levels  
> /  
> &NAMDPHY  
> NVEXTR=25,&nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! set number of tendency output fields (see table)  
> NCEXTR=137, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; ! edit to correctly set number of full model levels e.g. 60, 91, 137 etc  
> /&NAMPHYDS  
> NVEXTRAGB(1:25)=91,92,93,94,95,96,97,98,99,100,101,102,103,104,105,106,107,108,109,110,111,112,113,114,115,&nbsp; ! define GRIB codes for the tendency fields  
> /

I followed the steps on the [How+to+control+OpenIFS+output](https://confluence.ecmwf.int/display/OIFS/How+to+control+OpenIFS+output) page and took the advise regarding PEXTRA/LBUD23 on the forum into account. I think I set everything correctly, but maybe I overlooked a detail?

The model runs until step 2 by the way. There is output for ICMUA\<expid\>+000000. This is the ifs.stat file:

> 14:29:19 000000000 CNT3 &nbsp; &nbsp; &nbsp;-999 &nbsp; &nbsp; 4.172 &nbsp; &nbsp;4.172 &nbsp; &nbsp;5.629 &nbsp; &nbsp; &nbsp;0:00 &nbsp; &nbsp; &nbsp;0:00 0.00000000000000E+00 &nbsp; &nbsp; &nbsp; 0GB &nbsp; &nbsp; &nbsp; 0MB
> 
> &nbsp;14:29:42 A00000000 STEPO &nbsp; &nbsp; &nbsp; &nbsp;0 &nbsp; &nbsp;26.132 &nbsp; 26.132 &nbsp; 28.317 &nbsp; &nbsp; &nbsp;0:26 &nbsp; &nbsp; &nbsp;0:28 0.49588412029193E-04 &nbsp; &nbsp; &nbsp; 0GB &nbsp; &nbsp; &nbsp; 0MB
> 
> &nbsp;14:29:42 0AA000000 STEPO &nbsp; &nbsp; &nbsp; &nbsp;0 &nbsp; &nbsp; 0.000 &nbsp; &nbsp;0.000 &nbsp; &nbsp;0.000 &nbsp; &nbsp; &nbsp;0:26 &nbsp; &nbsp; &nbsp;0:28 0.49588412029193E-04 &nbsp; &nbsp; &nbsp; 0GB &nbsp; &nbsp; &nbsp; 0MB
> 
> &nbsp;14:29:53 FULLPOS-S DYNFPOS &nbsp; &nbsp; &nbsp;0 &nbsp; &nbsp;11.476 &nbsp; 11.476 &nbsp; 11.600 &nbsp; &nbsp; &nbsp;0:37 &nbsp; &nbsp; &nbsp;0:40 0.49588412029193E-04 &nbsp; &nbsp; &nbsp; 0GB &nbsp; &nbsp; &nbsp; 0MB
> 
> &nbsp;14:30:10 0AAA00AAA STEPO &nbsp; &nbsp; &nbsp; &nbsp;0 &nbsp; &nbsp;16.431 &nbsp; 16.431 &nbsp; 16.520 &nbsp; &nbsp; &nbsp;0:54 &nbsp; &nbsp; &nbsp;0:56 0.49588412029193E-04 &nbsp; &nbsp; &nbsp; 0GB &nbsp; &nbsp; &nbsp; 0MB
> 
> &nbsp;14:31:17 0AAA00AAA STEPO &nbsp; &nbsp; &nbsp; &nbsp;1 &nbsp; &nbsp;66.838 &nbsp; 66.838 &nbsp; 67.221 &nbsp; &nbsp; &nbsp;2:00 &nbsp; &nbsp; &nbsp;2:04 0.49394979039009E-04 &nbsp; &nbsp; &nbsp; 0GB &nbsp; &nbsp; &nbsp; 0MB
> 
> &nbsp;14:32:47 0AAA00AAA STEPO &nbsp; &nbsp; &nbsp; &nbsp;2 &nbsp; &nbsp;88.038 &nbsp; 88.038 &nbsp; 89.921 &nbsp; &nbsp; &nbsp;3:29 &nbsp; &nbsp; &nbsp;3:33 0.49186674044295E-04 &nbsp; &nbsp; 768GB &nbsp; &nbsp; &nbsp; 0MB

Does someone have advise? It just can be that I reached the limits of the HPC with the T1279 resolution.

Thanks in advance.

Great greetings,

~Thomas Batelaan

---

<div class="post-metadata">

**Author:** ![Marcus\_Koehler](https://forum.ecmwf.int/user_avatar/forum.ecmwf.int/marcus_koehler/32/1000_2.png) [@Marcus\_Koehler](https://forum.ecmwf.int/u/Marcus_Koehler)\
**Post date:** [28 April 2022 11:12 UTC](https://forum.ecmwf.int/t/openifs-lbud23-tendencies-out-of-mem/12175/2 "2022-04-28T11:12:20Z")

</div>

Hi [Thomas](https://confluence.ecmwf.int/display/~thomas.batelaan@wur.nl), I would first check the memory requirements of the job.&nbsp; You are running OpenIFS at 9 km resolution which is rather memory intensive, including PEXTRA fields multiplies the memory requirements in my experience.&nbsp; My suggestion would be to begin with a coarser grid and verify first that your experiments work at that resolution.

Cheers,&nbsp; Marcus

---

<div class="post-metadata">

**Author:** ![Glenn\_Carver](https://forum.ecmwf.int/letter_avatar_proxy/v4/letter/g/e495f1/32.png) [@Glenn\_Carver](https://forum.ecmwf.int/u/Glenn_Carver)\
**Post date:** [28 April 2022 11:15 UTC](https://forum.ecmwf.int/t/openifs-lbud23-tendencies-out-of-mem/12175/3 "2022-04-28T11:15:26Z")

</div>

Hi Thomas,

You don't give the actual error message reported by the job but I believe you that the job has run out of memory. T1279 is a very high resolution and adding all the pextra arrays for additional diagnostics will create alot more 3D fields.&nbsp; The other point to bear in mind is the amount of model output still will produce!

I assume you have already got as much memory allocated to the job as you can. Have you tried using additional nodes and underpopulating the nodes (reduce MPI tasks per node) to increase job memory?

If not, I'd suggest dropping down to a lower resolution, say T799 (which is still high) and making sure the job will work at a lower resolution first as Marcus suggests.

---

<div class="post-metadata">

**Author:** ![Thomas\_Batelaan](https://forum.ecmwf.int/letter_avatar_proxy/v4/letter/t/ee7513/32.png) [@Thomas\_Batelaan](https://forum.ecmwf.int/u/Thomas_Batelaan)\
**Post date:** [28 April 2022 14:10 UTC](https://forum.ecmwf.int/t/openifs-lbud23-tendencies-out-of-mem/12175/4 "2022-04-28T14:10:12Z")

</div>

Dear Ryan and Marcus,

Thanks a lot for your rapid replies.

I understand that T1279-jobs consumes a lot of memory. The fat-partition of our cluster has 128 cores per node and 1TiB memory per node/8GB memory per core – there are two partitions with even more memory but there you need special rights. If I interpret the error-file correctly I am really close to the limit of the cluster (see text-files with\_pextra and without\_pextra specs) in the without\_pextra job, so if it needs twice as much as what [Marcus Koehler](https://confluence.ecmwf.int/display/~damk) says it goes easily over it.

For the jobs I requested 384 cores (3 nodes with each 128 cores) with 1 core per task:

> #!/bin/bash
> 
> #SBATCH -t 02:30:00
> 
> #SBATCH --exclusive
> 
> #SBATCH --partition=fat
> 
> #SBATCH --ntasks=384
> 
> #SBATCH --cpus-per-task=1
> 
> #SBATCH --output=oifs.output%j.txt
> 
> #SBATCH --error=oifs.error%j.txt
> 
> ""
> 
> Setting the environment etc.
> 
> ""
> 
> srun master.exe

I am not sure how I can underpopulate more – I am still a beginner in HPC computing so do I understand correctly that I now use 128 MPI tasks&nbsp; per node because I requested 1 core per task on a cluster with 128 cores per node.

Thanks in advance,

Great greetings,

~Thomas

[without\_pextra.txt](https://forum.ecmwf.int/uploads/short-url/6mzY8SwT3S4QncXszkeILfr51Xa.txt) (3.7 KB)

[with\_pextra.txt](https://forum.ecmwf.int/uploads/short-url/8gXbgdKCWYSgRpZmIfTzIlo0ceZ.txt) (3.0 KB)

---

<div class="post-metadata">

**Author:** ![Glenn\_Carver](https://forum.ecmwf.int/letter_avatar_proxy/v4/letter/g/e495f1/32.png) [@Glenn\_Carver](https://forum.ecmwf.int/u/Glenn_Carver)\
**Post date:** [28 April 2022 14:33 UTC](https://forum.ecmwf.int/t/openifs-lbud23-tendencies-out-of-mem/12175/5 "2022-04-28T14:33:16Z")

</div>

Have you run this at a lower resolution than T1279?&nbsp; It's not a good idea to run at the highest resolution first until you know the job works correctly at lower resolutions.

If you have, then run the T1279 job without the pextra turned on first, to make sure it works and you can see the memory requirement without pextra. Underpopulating involves increasing the node count to get the required memory for the job (once you know what it is), and then reducing the mpi task count per node to be less than the available cpus per node.&nbsp; Your local HPC support should be able to help.

---

<div class="post-metadata">

**Author:** ![Thomas\_Batelaan](https://forum.ecmwf.int/letter_avatar_proxy/v4/letter/t/ee7513/32.png) [@Thomas\_Batelaan](https://forum.ecmwf.int/u/Thomas_Batelaan)\
**Post date:** [2 May 2022 09:23 UTC](https://forum.ecmwf.int/t/openifs-lbud23-tendencies-out-of-mem/12175/6 "2022-05-02T09:23:49Z")

</div>

Thanks all! With your help I managed to get it running (and also got more feeling about memory and so). Just requesting more nodes did the job.

(And yes, I already got it working with and without PEXTRA on a lower resolution before I scaled up)
