Mohamed Arafa, RHCE, RHOS-SA, Woodbadge, ITIL

Senior Site Reliability Engineer at SugarCRM

Cary, North Carolina, United States

About

Agile IT Infrastructure/Operations Management through DevOps Qualitative performance driven technologist with a unique combination of architecture, DevOps, training and management expertise leveraging 15+ years in financial and enterprise sectors. Focused on quality infrastructure, documentation, process, management and DevOps of the enterprise cloud sector. Specialties: OS: AIX SLES OpenSuSE CentOS Certified RHEL 5.x Fedora Linux Mandriva Linux RHEL6 RHEL7 Ubuntu LTS Hardware: IBM DS4800/3200 P570 HMC x3650 x3850 M2 BladeCenter S/E/H IBM Tape Library - LTO3853 HS21/22 HP DL580/380 Cisco_UCS HP_C7000 Storage: IBM Storage Manager RDAC IBM Director Software: NasdaqOMX X-Stream ORC's CameronFIX Zenoss Network Monitoring and Alerting DokuWiki MantisBT TSM Drupal IPCop LTSP Bash Scripting ISS Real Secure Suite Snort wireshark MySQL openfire XMPP server nessus sendsms3 zarafa dansguardian rpmbuild Oracle RAC OpenStack RoundCube Mail RDO asterisk RHN Satellite spacewalk ansible python jenkins docker kubernetes github Available to discuss new opportunities in the Cary, Raleigh-Durham, RTP area or virtual/remote positions. Contact me at [email protected]

Experience

  • Senior Site Reliability Engineer at SugarCRM
    Aug 2020 - Present · 6 yrs

  • Open Source Advocate at Free and Open Source Software Movement
    Jul 1995 - Present · 31 yrs 1 mo

    • Advocated Open Source whenever the chance presented itself. • Talked about the advantages of open source substitute software • Tested and improved FOSS software by personally using it and reporting bugs when necessary and following through to ensure the developer is given proper information to fix the bug. • July 2014 - Gave a presentation to the Government of Egypt and industry entitled "Selling OpenStack to Egypt", an audience of 60 people

  • IBM (4 yrs 11 mos)
    • Site Reliability Engineer
      Oct 2015 - Aug 2020 · 4 yrs 11 mos

      • automated the backup of the etcd db to s3 storage • fine tuned and gained significant speed increases by tweaking our ansible playbooks • implemented a monthly jenkins job to run our deployment playbooks against all our servers. since our security fixes go in to our playbooks, the servers need to be updated on a regular basis • setup the process and migrated some jenkins jobs to use jenkins job builder and documented the process for others to replicate • migrated a set of monitoring services from an old deployment pipeline to a newer pipeline • worked with other squads to deploy brand new kubernetes clusters • set up ansible tower and RHN Satellite as POC • wrote a pipeline to build and distribute pre made and secured VM images to all our world wide data centers • acted as unofficial agile project manager for a couple of projects, setting schedules and weekly targets • led scrum calls, represented the squad in managerial weekly update meetings

    • Software Engineer/IBM Kubernetes SRE
      Oct 2015 - Jun 2020 · 4 yrs 9 mos

      As a DevOps Site Reliability Engineer/SRE, supported the IBM Kubernetes Service hardware and software infrastructure and occasionally act as level 3 kubernetes / openshift on kubernetes support. Used Continuous Integration/Continuous Development (CI/CD) methodologies to automate tasks using the available toolsets available to me. Toolsets have included: docker containers, openstack, kubernetes, jenkins, github, travis, ansible Projects have included: • planned and developed a Jenkins pipeline to build a gold image for IBM’s fleet of kubernetes workers that used s3, Jenkinsfiles, JJB, bash, python, slack, patching, ansible and more • managed the migration of users, projects and jobs from one jenkins instance to another • managed, planned, communicated and implemented the upgrade of docker daemons on the infrastructure servers through the promotion process. • used widely available knowledge and process to automate the scheduling and identification of nefarious users against our TOS, disabling their containers, and automated reporting to the responsible project office to communicate with clients and disable their accounts. • participated in a project to continuously improve the use of ansible to customise our post-install playbooks. Including updating these playbooks for security vulnerabilities as they show up. • member of a small team that mitigates vulnerabilities identified by nessus. Interacts with the security group to ensure vulnerabilities stay mitigated. • Wrote and used ansible playbooks, shell scripts, etc to run tasks and mitigate issues as they occur. • Setup an apt-repo and integrated it into our deployment process via our ansible playbooks • documented and automated the set up of the infrastructure environment of our clusters • on a personal basis, explored the AWS experience via a private account

  • Customer Support Engineer, virtual Managed Services at Cisco
    May 2015 - Oct 2015 · 6 mos

    Supported Cisco's virtual Managed Services orchestrated by tail-f's confd

  • Cloud Escalation Engineer at Verizon Enterprise Solutions
    Nov 2012 - Apr 2015 · 2 yrs 6 mos

    • Responsible for the health of over 1 million VMs hosted on 10,000 ESX hosts running on over 9500 UCS and HP blade servers connected to over 200 SANs with 70 Petabytes of data, in addition to thousands of dedicated VMware, Windows and Linux servers on HP Proliant Server hardware • As a member of the team, one of my responsibilities was not only to take ticket escalations from Tier1 but also to reduce ticket count by optimising and enhancing processes, hunt bugs and follow through to a resolution either by identifying the source of the bug or forwarding it to the developers if needed. • Another responsibility was training others, and as a byproduct of that aim I wrote and recorded 5 VZLearn modules that at last count certified 90 people across several disciplines to work on ECME. • Added to the repository of knowledge by either adding new documents or enhancing existing ones leading me to become 2013's largest contributor to the Knowledge Base. • Audited and discovered terrabytes of unbilled storage that I turned into billable income • Change Advisory Board member. • Used HP Operations Orchestration to push out scripts I wrote to hundreds of managed servers for maintenance reasons