Coburg North, Victoria, Australia
Started life professionally as a telescope operator who did software development, systems integration and systems administration during my literal day job (and sometimes at night. No testing like testing in prod). Moved onto the Bureau of Meteorology, where I was originally hired as a System Administrator in the Midrange Platforms group, and was responsible for, amongst many other systems, one of the busiest (and most responsive) websites in Australia. Got seconded as a Senior Systems Administrator in the High Performance Computing group, where I successfully stood up, integrated and operated the current supercomputers that have enabled three major generational changes to the Bureau's forecast and climate models, while building out the next hardware upgrade and planning the upgrade beyond that. Worked at CMD @ Mantel Group in the first half of 2023 to explore infrastructure as code and cloud computing in AWS, and La Trobe University in 2024, documenting and improving runbooks and advocating for more of a DevOps culture between our staff and stakeholders. Interests are physically interacting with heavy (800tonnes in one example) bits of engineering at a low level, or immersing myself in data. I'd probably enjoy being an embedded programmer, and have way too many embedded microcontrollers controlling bits of hardware at home.
Physical rack and stack of DDN hardware, with subsequent install, configuration, tuning, integration and troubleshooting of the storage software stack from NFS and Lustre etc clients through to the kernel. Solution documentation, technical implementation plans, customer training, and contributions to hardware and software manuals. Manage customer relationship during installation and with subsequent managed services.
Through close working with stakeholders, fostered a culture of shared responsibility. Collaborated in brainstorming preliminary design proposals. Found solutions and paths forward for clients stuck on legacy platforms. Guided team members on modern configuration and infrastructure management practices.
Documented existing infrastructure and practices, identified deficiencies, improved existing runbooks, collaborated in reducing out-of-hours callouts. Built software for our researchers, and advocated for more of a devops culture between our staff and stakeholders.
Identified client platform deficiencies, provisioned new infrastructure. Fixed faulty monitoring resulting in excessive callouts and inadequate alerting of critical resources, and improved customer runbooks. Dealt with backups and patching, cost optimisation and compliance, Infrastructure as Code (IaC) drift detection and remediation. Mentored junior consultants and performed stakeholder engagement and customer handoff. Specialties: ---------------- AWS, Grafana, Terraform, CDK
Seconded to the newly constructed Scientific Computing Services (High Performance Computing) group to stand up, harden, integrate, operate and continually improve the Development and Highly-Available Operational Cray XC40 supercomputers and associated support systems, distributed storage (Lustre, GPFS, NFS) and postprocessing clusters, whilst the security posture of the organisation was rapidly increasing. Managing observability at scale, developed self-healing functionality and compute and storage failover procedures to ensure > 99.95% uptime of the operational petascale platform, maximising utilisation of all resources, and minimising out-of-hours callouts. Trained colleagues in development tools and practices such as git, code reviews and infrastructure as code (IaC). Developed puppet and Ansible roles and playbooks to automate the provisioning/maintenance of supporting infrastructure in the development, staging and production realms. Built and operationalised the replacement Ansible-based XC50/CS500 systems, while planning the DR systems beyond that. Specialties: ---------------- Hardware, clusters, distributed systems, VMware, KVM, Corosync/pacemaker, HA, Linux, backups, monitoring, scripting, Bash, Perl, HPC, puppet, Ansible, F5, Site Reliability Engineering
System Administrator in the Midrange Platforms group. Managed external dual-sited load balanced WWW clusters (serving ∼100 million hits per quiet day, with > 10× peak demands). Helped design and manage FTP, DB and app clusters and DNS, LDAP etc servers along with disaster recovery services; supported the critical Highly-Available services (Central Message Switching Service and others), and the operational, development & testing environments for core and other internal Unix services (700 servers under our control, I was primary contact for 290 servers). Managed RHEL on VMWare as well as physical clustered hardware, using Veritas Clustering Server, Pacemaker/Corosync, VMWare and others for high availablility. Problem, security and incident response included recovering faulty clusters and corrupted filesystems with negligible downtime well within SLO, liaising with upstream vendors and contractors to replace faulty hardware, debugging kernel issues causing repeated crashes or cache incoherency, and emergency security mitigation while waiting for official upstream patches. Liaised with other business units to solve complex cross discipline issues from the core network to application issues. Migrated VMs and critical services between datacentres without downtime by combination of lift-and-shift and geographically spanning cluster heartbeats, storage and VLANs in datacentre decommissioning projects. Tech lead in splitting out and virtualising shared HA clusters including front and backend for ftp.bom.gov.au, reg.bom.gov.au and the website (migrating between incompatible storage technologies with a large data change rate with no allowable downtime). Sat on evaluation panel for next generation of machines to replace core infrastructure, saving many hundreds of thousands of dollars over what had been previously considered. Specialties: --------------- Hardware, clusters, distributed systems, VMware, Veritas VCS, HA, Linux, backups, monitoring, scripting, puppet, tripwire.