Infrastructure

Hardware, software and facilities

The laboratory is being built in two stages. What is listed as installed is approved and in service; what is listed as planned is scheduled and not yet operational.

Stage one

Installed and approved

Lenovo ThinkStation P8

Primary research workstation

Processor
AMD Ryzen Threadripper PRO 9975WX — 32 cores, 64 threads
GPU
NVIDIA RTX PRO 6000 Blackwell Workstation Edition, 96 GB
Memory
256 GB DDR5-6400 RDIMM with error correction
Storage
4 TB PCIe Gen5 NVMe
Operating system
Windows 11 Pro

Deep learning workstation

Model training and development

Processor
Intel Core Ultra 7 270K
GPU
NVIDIA GeForce RTX 5090, 32 GB GDDR7, 512-bit
Memory
128 GB DDR5-5600
Storage
6 TB PCIe Gen5 NVMe (2 TB + 4 TB)
Cooling
Noctua NH-D15 G2 air tower

What the specification is for

Ninety-six gigabytes of error-correcting GPU memory on the ThinkStation puts frontier-scale open-weight models on institutional hardware, which removes per-query API cost and keeps text under confidentiality obligations inside the building. Models of roughly 120 billion parameters run locally at reduced precision.

Thirty-two cores and 256 GB of system memory carry the parallel simulation work behind the laboratory's methods papers. Error correction throughout protects the numerical integrity of published estimates, which matters when a Monte Carlo design runs for days.

Why two machines

The two are not interchangeable. The RTX 5090 workstation is the faster card per dollar and handles training runs, image pipelines and day-to-day development, but its 32 GB of memory caps the size of model it can hold.

The ThinkStation exists for the work that cap rules out: large-model inference, high-memory panels and long simulation designs where error-correcting memory matters. Running development on one and the memory-bound work on the other keeps both busy.

The constraint this removes

Shared campus systems limit researchers to 64 GB of memory. Several laboratory datasets, including a 170-million-row signal panel, cannot be held in that budget at all.

Stage two

Planned for 2027–2028

A dedicated laboratory space with additional compute, storage and collaborative facilities. None of the following is in service yet.

Laboratory space

  • Rooms 6429 and 6430, Building 6B Level 4
  • 47.3 m² of purpose-fitted research space
  • Viewing room and focus discussion room

Additional compute

  • Dell Precision 7875 Tower workstation
  • Custom high-memory build with H100-class GPU
  • Network-attached storage for datasets and backup
  • Two compact desktops with dual 27-inch monitors

Collaboration

  • ViewSonic 110-inch 4K touch interactive panel
  • Large-format display for meetings and briefings
  • Weekly working sessions and annual briefings

Delivery schedule

  1. Q3 2027
    Design and build consultant engaged
  2. Q4 2027
    Construction begins
  3. Q1 2028
    Construction complete
  4. Q2 2028
    Laboratory fully operational
Stack

Software and environment

Statistical computing

  • R, with the econometrics and spatial stacks
  • Python for machine learning and text pipelines
  • Stata for partner workflows under migration
  • SLURM scheduling for batch and array jobs

Local model inference

  • Open-weight models in the Llama, Qwen and Mistral families
  • Fine-tuning on laboratory data without external transfer
  • Batch inference for high-volume benchmark studies
  • No third-party API keys in the research path

Reproducibility

  • Version-controlled analysis code
  • Replication packages released with published work
  • Fixed random seeds and recorded environment state
  • Separation of raw data, derived data and outputs
Governance

Why the compute is on site

Several laboratory programmes work with data whose ethics approvals, participant agreements or partner contracts restrict where processing may happen. On-premise infrastructure satisfies those conditions without an exemption.

Custody stays local

Clinical signal records, licensed corpora and partner datasets are held and computed on machines inside the building, under Monash custody throughout.

Access is mediated

Analysts submit code against restricted datasets and receive aggregate output. Each query is recorded, so a data custodian can establish who used a dataset and for what.

Backup is encrypted

Research data is replicated and encrypted at rest, so that the loss of any single machine does not put a project's work or its recovery at risk.

Access

Using laboratory compute

Colleagues across the School and the campus can run work here when a project exceeds what a standard machine holds, or when its approvals rule out a commercial cloud service.

External organisations can arrange hosted compute and model inference through a partnership. Tell us the size of the data, the method, and any conditions attached to the data, and we will tell you whether it fits.