Rachel Lee

Site Reliability Engineer @ HDB | Observability | AIOps

Singapore, Singapore

About

Site Reliability engineer focused on automation and observability at scale — handling Splunk-driven monitoring across 1,000+ assets. Passionate about building reliable, self-service infrastructure through code.

Experience

  • Site Reliability Engineer at Housing & Development Board
    Jul 2024 - Present · 2 yrs 1 mo

    - Spearheading an AIOps observability initiative at HDB using Splunk IT Service Intelligence to improve monitoring, logging, and system visibility across infrastructure - Handled integrations between Splunk and third-party tools including ServiceNow, SolarWinds, Dynatrace, and Pingdom to enable centralized monitoring and alerting - Oversaw platform management and configuration for a Splunk environment covering 1,000+ assets — 800+ via Splunk forwarders/agents and the remainder via syslog collection from hardware devices (Checkpoint, Palo Alto, F5) - Identified key SLIs and implemented appropriate thresholds, leveraging Splunk's machine learning capabilities for predictive analytics and adaptive thresholding - Deployed SC4SNMP with Docker to collect SNMP data from hardware devices, including performance metrics (CPU, memory) and node status polling - Collaborated with multiple stakeholders to verify integrations, manage data onboarding, and handle other BAU (business-as-usual) operations

  • Infrastructure Engineer at DBS Bank
    Sep 2021 - May 2024 · 2 yrs 9 mos

    - Developed a Python-based orchestration tool using vCenter API to automate end-to-end VM provisioning, reducing manual setup time by 30% - Designed templating workflows using cloud-init and Jinja to standardize and accelerate provisioning for microservices architecture - Improved system observability by extending the Elastic stack, giving teams better visibility into performance and errors - Created bash scripts to automate configuration file changes in Linux environments, reducing chances for human error - Built IaC pipelines using Jenkins and Packer to automate image provisioning for vCenter environments - Led migration projects transitioning from local filesystem to cloud storage using in-house tools, cutting mount dependencies by 50% - Drove cross-functional collaboration — leading scope-gathering sessions with support and application teams to resolve blockers and keep delivery on track - Performed functional and performance testing and troubleshooting with various stakeholders to ensure application reliability