FAQ¶
Common questions about using DelftBlue, grouped by topic. Use the contents to jump to a question, or your browser's search.
About DelftBlue & getting started¶
I found an error or outdated information in the documentation, or would like to contribute¶
The documentation is maintained by a very small number of people, so we are very
happy if you want to contribute in any form. Open an issue
on this GitLab page, selecting
either the update_page or new_content template. If you do not use one of the
templates, make sure to assign the issue to a DHPC employee, because otherwise no
one is notified.
Is DelftBlue the right machine for me?¶
It depends on what exactly you are trying to do. DelftBlue is a supercomputer, intended to allow scaling up applications. It may require a significant effort to get started if you are not used to working on HPC systems.
TU Delft also provides a range of different services, including personal workstations, physical or virtual machine servers, cloud services, faculty based local servers, and many others. If you are looking for an IT solution, but are not sure what exactly is the best option for you, talk to your faculty's IT manager.
I am new to Linux. Help!¶
The basics of working with Linux and remote systems can be relatively easily self-learned with online materials. Start with checking out the relevant Software Carpentry courses:
- Linux command-line basics
- Linux command line (more advanced material)
- Introduction to High-Performance Computing
We also prepared a super-short self-learning page specific to DelftBlue for novices: Crash-course on DelftBlue for absolute beginners.
Where can I find courses on using and programming HPC systems like DelftBlue?¶
- DCSE courses are generally run on DelftBlue. They have a logical build-up to get you "from zero to hero", and are accepted as disciplinary training by the graduate schools.
- SURF courses, in particular the "Introduction to Supercomputing" series, are geared toward the national computing infrastructure, but it is reasonably similar to DelftBlue.
- The e-Science Center offers various courses geared towards scientists, although not necessarily geared towards HPC.
Where do I get help?¶
For technical issues, please use the TU Delft Self-Service Portal DelftBlue category under "Research support".
For discussions about specific software, compilers, etc., as well as exchanging knowledge with other users, please use the DHPC Mattermost forum.
I am unhappy with my disk quota, job time limitations, the number of GPUs I can use, etc.¶
We are implementing these policies not in order to annoy users, but in order to be (and remain) useful for as many users as possible. The policies are not set in stone, though, and your feedback is very welcome so that we can keep improving the system.
For more information on why we made these policies, see the DHPC Policies page.
Access & login¶
How do I get access?¶
By default, every employee with a <netid> is able to log in to DelftBlue and run
jobs as a guest. However, your jobs will run with relatively low priority, and if
the machine is busy, it may take a long time until they get scheduled.
To be able to run jobs with high priority, you can request access to your faculty's share.
To get full access, please use the form "Request Access" in the self-service portal.
I can not log in¶
Are you a student? If yes, you have to explicitly request access here.
Are you trying to access the right machine? DelftBlue is TU Delft's largest supercomputer; however, individual faculties and departments might still operate their own local servers and clusters. For example, the INSY department operates their own machine, known as DAIC. Make sure you are trying to connect to the right system!
Check the correct address of DelftBlue (login.delftblue.tudelft.nl), the correct
port (22), and the correct username (your <netid>). If you are using the SSH
config file, check that it has been set up correctly.
When you cannot make a connection (no prompt for username or password at all), check the Remote access to DelftBlue page for the ways to access DelftBlue from the outside world. If you are outside of campus, use EduVPN!
When you receive an authentication failure when you enter your password, check if
you are using the correct <netid> and password. For example, try to log in to
the TU Delft webmail.
Finally, if all the above points are resolved, check the status of DelftBlue on the web page of the Delft High Performance Computing Centre. Make sure the system is up, and there are no interruptions or planned maintenance.
I can log in, but I don't have a /home folder¶
Your /home (and /scratch) directory is created on your first shell login
to DelftBlue. If your very first login was through Open OnDemand instead, you get
a "Home directory not found" screen — click Open Shell on it, type logout in
the shell that opens, then click Reload Webportal. See
Open OnDemand → Accessing Open OnDemand
for the screenshot and context.
I can log in, but I cannot submit a job¶
You should be able to submit jobs to the queue using the sbatch command (see
the Slurm scheduler page). If there are any specific errors
in your submission script, those should either be displayed in your terminal
window, or in the slurm-XXX.out file.
Please note that the Open OnDemand web interface does not have a working job
template at the moment! If you use the web interface, make sure to copy the
correct #SBATCH commands from one of the
examples in the documentation!
Jobs & scheduling¶
How exactly is the priority for my job calculated?¶
This question is hard to answer. Roughly, the priority is calculated based on your previous usage, the overall usage of the account (i.e. of all users in the research/education/project account you use), and the waiting time of the job so far. However, priority is not the only aspect determining when your job will run: for example, short-running jobs that can be spread over cores on multiple nodes may get ahead by so-called "backfill scheduling" without pushing back other jobs.
The configuration of the Slurm scheduler on DelftBlue is maintained by our
experienced HPC admin team. It is completely transparent, and if you want to know
details, you can use the command scontrol show config on DelftBlue to get the
full list of settings: at the time of writing, for example, there are 17
parameters containing Priority, just to illustrate why we do not give you an
"exact" answer. Information on what these parameters mean can be found in the
Slurm scheduling configuration guide.
Why does my job not run?¶
Every user with a <netid> is able to log in to DelftBlue. However, your jobs
will run with rather low priority in the innovation account, and if the machine
is busy, you may have to wait for several days. You can request access to your
faculty's research share via
this form
in the self-service portal.
If you are a registered user and your jobs still do not get scheduled within a day
or so, the column REASON in the output of squeue --me will give you a clue.
The Resources you request may not be available at the moment, or they are
reserved for users with higher Priority (see
How exactly is the priority for my job calculated?
above). Note that the resources you request are determined by the job script and
default settings for options you omit, and may not always be obvious to you: for
example, if you specify --ntasks=1 --cpus-per-task=1 --mem-per-cpu=100GB, this
job would occupy about half a node (24 cores in Phase 1) because of the high
memory usage, even if only one core (2 percent of the CPUs) is being requested.
Finally, a personal or group limit may apply. These are implemented to ensure fair scheduling of jobs. For details see the Slurm documentation on job reason codes.
Can I get an estimate when my job will run?¶
If your job is already in the queue and pending, you can use squeue --start to
get a rough estimate. This will only output a predicted starting time if the job
is already waiting for resources.
You can also get a prediction before submitting the job, for example to evaluate
which partition to use or how much memory to ask per CPU to get the job started
sooner. The command sbatch --test-only <job script> prints the estimated start
time (based on all information available at that moment) without actually
submitting the job.
Can I log in to the node(s) on which my job is running?¶
Sometimes you may want to check how your job is running using commands like top,
nvidia-smi, etc. In order to get a console "inside your Slurm job", first find
the job's ID using
and then start an interactive bash session as follows (here for 30 minutes):
See the Slurm interactive jobs page for details on interactive jobs.
Slurm fails with out-of-memory (OOM) error¶
If your job is killed with slurmstepd: error: Detected ... oom-kill event(s), it
used more memory than was requested. Raise --mem-per-cpu in the submission
script to a realistic value — see
Trouble-shooting your job → Slurm out-of-memory (OOM) error.
Software & environment¶
I can not find and/or load software modules¶
We provide two versions of the software stack; typically you should pick the latest one, 2026 in this example:
[<netid>@login03 ~]$ module avail
------------------------------------- /apps/noarch/modulefiles -------------------------------------
2025
2026
To make the contents of the software stack 2026 available, use:
This system is based on lmod.
The module organisation is hierarchical. This means that the modules you see
depend on the ones you loaded: in particular, if you don't load openmpi you
don't see modules like hdf5 that depend on it. To find modules in the hierarchy,
you can use the module spider command.
Check out the DHPC modules page for more info.
After loading certain modules, the nano (editor) command does not work any more¶
This is caused by incompatible libraries in the search path. You can avoid it by
putting the following line in your .bashrc file, and running exec bash to
refresh the active session:
Locale warnings when I log in or start a GUI app¶
The DelftBlue nodes use LANG=C.UTF-8 by default, which works fine. You only see
messages like
locale: Cannot set LC_CTYPE to default locale: No such file or directory
bash: warning: setlocale: LC_ALL: cannot change locale (en_US.UTF-8)
Detected locale "C" ... Qt ... has switched to "C.UTF-8" instead
when your session requests a locale that is not installed — only C,
C.utf8 and POSIX are. The usual sources:
- Your terminal forwarding
LANG/LC_*over SSH. macOS Terminal and iTerm2 do this by default (in iTerm2: Settings → Profiles → Terminal → Set locale environment variables automatically), typically sendingLANG=en_US.UTF-8. - A
LANG/LC_*line in your~/.bashrcor~/.profileon DelftBlue. - The
abaqus/2024module, which setsLC_ALL=en_US.UTF-8itself — you get the warning onmodule load abaqus/2024.
For almost every tool these messages are harmless — the tool falls back to C
and keeps running. To get rid of them, stop your terminal from forwarding the
locale, or force one that exists for the session:
Use C.utf8, not bare C or POSIX (those are ASCII-only and can mangle
non-ASCII output). If you set it before module load abaqus/..., note that the
abaqus module overrides LC_ALL when it loads, so re-run the export afterwards.
I compiled my code, but it fails with Illegal Instruction errors on some nodes¶
Please be aware that most nodes have Intel CPUs, except for gpu-v100 nodes, which
use AMD CPUs (see the DHPC Hardware page for
details). Depending on your application, you might need to compile two different
versions of your program to run it on the respective nodes.
"Phase-2" nodes have a later generation of Intel CPUs, which may also cause code to be generated that does not run on the older hardware.
Please also note that available software modules for compute and GPU nodes
might be different for this reason.
I have difficulties with compiling a specific software package¶
Many software packages are already available on DelftBlue as software modules. Instructions for certain frequently used packages can be found in the Software recipes (Howtos) section (see the navigation menu).
Intel compiler takes forever or does not work at all¶
TU Delft's licence for the Intel compiler suite allows for 5 users to compile
their software simultaneously. If your ifort compilation takes longer than
expected, or does not start at all, try again in a bit. If it still does not work,
contact us.
See the Intel compilers howto for more information.
My Intel-compiled code runs normally on the login node, but is super slow upon submission to the queue¶
For the Intel MPI to be able to operate correctly, it must be configured to work
together with Slurm. Otherwise the Intel MPI may bind all requested threads to the
same CPU. To configure the Intel MPI to work with Slurm, set Intel MPI to use the
Slurm PMI interface, and use srun instead of mpirun.
Use the following in your submission script:
FIXME: Should this be /usr/lib64/libpmix.so or /usr/lib64/libpmi2.so or
/usr/lib64/libpmi.so? (If you find out, please report it back to the DHPC team.)
Then invoke your binary with srun. Typically there is no need to specify
arguments as they are taken from the #SBATCH commands in the job script by
default.
Storage & data¶
I cannot access my files on the TU Delft home, bulk, umbrella drives¶
On the login and transfer nodes, these drives are mounted under /tudelft.net. On
all other nodes, they are not directly accessible. If you get the error message
"permission denied" when accessing your directories under /tudelft.net on the
login or transfer nodes, this is likely caused by an expired Kerberos ticket. You
can refresh your Kerberos ticket by typing kinit and entering your password upon
request. Details can be found on the
Data transfer to DelftBlue page.
I would like to ensure that I remove my data from the local disk when the job ends¶
Data in the local /tmp storage on a node is not removed automatically when a job
ends. The job needs to explicitly remove any data that it stores there. If it
doesn't, other jobs won't be able to use that space and may fail.
When a job times out or is cancelled, the job script stops execution completely, and any cleanup commands at the end of the script won't run. Instead, use the code in the example job script below to automatically clean up on exit or cancellation.
#!/bin/bash
# Your sbatch specifications go at the top
#SBATCH ...
# Specify and create a unique local temporary folder location
tmpdir="/tmp/${USER}/${SLURM_JOBID}"
/usr/bin/mkdir --parents "$tmpdir" && echo "$tmpdir created successfully."
# Clean up the local folder when the job exits (even when it times out or is cancelled)
function clean_up {
/usr/bin/rm --recursive --force "$tmpdir" && echo "Clean up of $tmpdir completed successfully."
exit
}
trap 'clean_up' EXIT
# Your own script goes below
# You need to explicitly specify and use this "$tmpdir" location in your code!
# (Below are some example lines, uncomment and modify if you want to use them.)
#/usr/bin/cp "/scratch/${USER}/some_file" "${tmpdir}/some_file"
#srun <your command> ... <option to specify temporary folder> "$tmpdir" ...
#/usr/bin/mv "${tmpdir}/result_file" "${HOME}/result_file_run123"