Site Reliability Engineer

Full Time
Cork, County Cork
Posted
Job description
Company Description Since 2004, Mandiant has been a trusted partner to security-conscious organizations. Effective security is based on the right combination of expertise, intelligence, and adaptive technology, and the Mandiant Advantage SaaS platform scales decades of frontline experience and industry-leading threat intelligence to deliver a range of dynamic cyber defense solutions. Mandiant’s approach helps organizations develop more effective and efficient cyber security programs and instills confidence in their readiness to defend against and respond to cyber threats.
Job Description


Mandiant is seeking a Fixed Term Contract Site Reliability Engineer to join our Production Operations SRE team. The SRE role will focus on elevating application and service performance and availability in support of our organization’s fast-evolving enterprise technology needs.

The SRE role actively targets risk to service availability for employees and customers by partnering with Engineering and Operations teams leveraging modern observability tooling and service restoration methodologies focused on automation and infrastructure as code where possible.

The ideal candidate will have a broad background spanning both applications and infrastructure. They will have direct experience in coding languages and core SRE practices and methodologies.

What you will do:

  • Monitor, measure and improve the reliability, availability and scalability of IT Infrastructure, applications and services
  • Identify manual routine operational practices and build robust automation capabilities using code and modern tools
  • Collaborate with Product Developers and business stakeholders to gather requirements for enabling and improving performance monitoring for applications and services
  • Engage in Incident response and participate in post-mortem analysis to investigate root cause and capture contributing factors for remediation
  • Perform analytics on previous incidents and trend/usage patterns to better predict issues and take proactive actions
  • Design and build custom tools as needed to support process optimization, challenging the status-quo and improving operational efficiency
  • Participate in 24*7 rotational shifts for handling production operation issues
  • Engage in service capacity planning and demand forecasting, software performance analysis and system tuning
  • Create meaningful dashboards/reports for application telemetry and infrastructure health for pro-actively identifying performance constraints and bottlenecks

Qualifications
  • Strong sense of ownership and an ability to drive cross-functional process improvement
  • Possesses excellent inter-personal written and verbal communications skills
  • Analytical and logical approach to problem-solving and a willingness to automate repetitive tasks and reduce manual/reactive workload
  • Perform logical and systematic search for the source of a problem in order to solve it and make the product or process operational again
  • Apply in-depth troubleshooting and debugging skills of systems, databases, and applications to get to root cause of the issue
  • Experience in administration/build/management of Linux or Windows systems
  • Foundational understanding of Infrastructure and Platform Technology stacks
  • Working knowledge of Infrastructure and Application monitoring platforms such as Logic Monitor, OpsGenie, Datadog, New Relic, Elastic / ELK Stack, Splunk, Thousand Eyes, CloudWatch, Big Panda
  • Experience participating in Major Incidents and Problem Management

Additional Information
  • Strong understanding of cloud-based architecture and cloud operations. Hands-on experience with public cloud technology
  • Service availability oriented mindset with a pro-active approach to problem solving. An ideal candidate should be able to develop automated solutions to prevent recurring problems
  • Strong understanding of Networking concepts and theories, such as different protocols (TCP/IP, UDP, routing protocols, etc), VLAN configuration, DNS, OSI layers, and load balancing
  • Understanding of security architecture and certificate management
  • Understanding of the core DevOps practices (CI/CD pipeline, release management etc.)
  • Ability to write code using any one modern programming language (Python, JavaScript, Ruby etc.). Additional scripting skills are preferred
  • Configuration management platform understanding and experience (Chef/Puppet/Ansible)
  • Prior experience in Cloud management automation tools (Terraform/CloudFormation etc.) is preferred
  • Experience with source code management software and API automation is preferred
  • Possesses the ability and willingness to challenge the status-quo and optimize current procedures and processes

Please note that this is a 12 month Fixed Term Contract.

seankuhnke.com is the go-to platform for job seekers looking for the best job postings from around the web. With a focus on quality, the platform guarantees that all job postings are from reliable sources and are up-to-date. It also offers a variety of tools to help users find the perfect job for them, such as searching by location and filtering by industry. Furthermore, seankuhnke.com provides helpful resources like resume tips and career advice to give job seekers an edge in their search. With its commitment to quality and user-friendliness, seankuhnke.com is the ideal place to find your next job.

Intrested in this job?

Related Jobs

All Related Listed jobs