Site Reliability Engineer
impact.com
About Impact.com
impact.com is the world’s leading commerce partnership marketing platform, transforming the way businesses grow by enabling them to discover, manage, and scale partnerships across the entire customer journey. From affiliates and influencers to content publishers, brand ambassadors, and customer advocates, impact.com empowers brands to drive trusted, performance-based growth through authentic relationships. Its award-winning products—
Performance (affiliate),
Creator (influencer), and Advocate (customer referral)—unify every type of partner into one integrated platform. As consumers increasingly rely on recommendations from people and communities they trust, impact.com helps brands show up where it matters most. Today, over 5,000 global brands, including Walmart, Uber, Shopify, Lenovo, L’Oréal, and Fanatics, rely on impact.com to power more than 225,000 partnerships that deliver measurable business results.
Your Role at impact.com
As the Site Reliability Engineer for the Content Intelligence & Regulatory Apps Group, you will own and grow our reliability practice for the systems that index, monitor and enrich social and web content on the impact.com platform. This role focuses on building and professionalizing rather than firefighting. Since the platform is stable, your mission is to implement SRE engineering disciplines (service-level objectives, observability, runbooks, and structured root-cause analysis) to ensure reliability is measurable, repeatable, and owned.
You will work with Squad leads, Platform Engineering and Cloud Operations and will own the SLO and RCA practice for the group while partnering closely with other internal and external squads. Your job is to provide them with the framework, tooling, and habits needed to run reliable services, and to act as the central point of contact for reliability across those teams. This is a software-engineering-led SRE role where you will read and write Java, instrument Spring services, tune the JVM, help harden infrastructure, batch processes, data and orchestration flows.
Our guiding principle is to prioritize system stability and data integrity above all else. Because these are business-critical systems, security and compliance are part of the reliability mandate, not an afterthought: you will build observability, audit trails, and operational practices that are secure and auditable by default, working alongside the central security and DevOps teams. Reliability, resilience, data integrity, and compliance take precedence over short-term feature velocity. This is a strong opportunity for an engineer ready to step into ownership and grow the role and themselves over time, with the support of the group.
What You'll Do
- Own the SLO/SLI practice for CIRA: Define meaningful service-level objectives and indicators for services on GCP with each squad. Establish error budgets and necessary baselines, applying extra rigor to flows that affect critical business processes.
- Build in security and compliance by default. Treat security and auditability as reliability properties: ensure critical business processes and data flows have the necessary audit trails, and if applicable, forensic-replay history needed for transactional-correctness and compliance obligations (e.g. SOX, and PCI-adjacent concerns). Champion secrets hygiene (HashiCorp Vault, GCP Secret Manager, SOPS), least-privilege access to production and data, and audited break-glass procedures. Partner with the Cloud Security and Cloud Platform teams rather than duplicating their function.
- Manage vulnerability and patch posture for services: Track and drive remediation of vulnerabilities across the JVM, Spring/Spring Boot dependencies, and container images; help establish patching expectations and surface security-relevant findings from quality gates (SonarQube) and secret scanning (ggshield) so they get prioritized alongside reliability work.
- Own and run root-cause analysis: Establish a consistent, blameless RCA practice for the group. Drive investigations toward durable fixes and preventative actions while partnering with the owning squad rather than working in isolation. Over time, improve the squads' own troubleshooting and post-incident habits.
- Build and mature observability for critical services: Become the group's point person for the observability stack (e.g., Grafana). Build and standardize monitoring, dashboards, tracing, and alerting using Open Telemetry that surface the health of systems, services, infrastructure and business/transactional processes — without leaking PII data into logs, traces, or dashboards — while creating reusable patterns that squads can adopt.
- Keep performance and resource metrics within thresholds: Track response latency, JVM heap/GC behavior, thread-pool saturation, CPU/memory consumption, cloud costs, error rates, and uptime across all services (e.g. Java, Spring Boot, NodeJs etc). Turn findings into prioritized improvements in collaboration with the squads.
- Improve batch and orchestration reliability: Help harden batch jobs and support the migration toward durable workflow orchestration to reduce partial-failure windows and manual idempotency.
- Support database performance and query optimization: Monitor the groups databases (e.g. MySql, SingleStore, Elastic, GCP Spanner) for slow queries, indexing, lock contention, and connection-pool health. Flag operational or performance risks in Liquibase migrations before they ship.
- Understand client-generated workloads: Learn how brand, agency, and partner workloads translate into resource consumption. Help confirm that consumption aligns with contract tiers and surface anomalous traffic patterns as both a reliability and a security signal, escalating suspected abuse to the security team.
- Establish alerting and runbook standards: Define the group's conventions for actionable alerts and runbooks to ensure on-call engineers across squads can respond quickly and consistently.
- Troubleshoot across the stack: Address issues in the JVM/application, Spring container, MySQL, Pub/Sub messaging, Feign/REST and legacy RMI/ integrations, GKE/GCE, network, and client-generated workloads. This requires growing depth in JVM and database performance (profiling and query optimization). Remediation may include code/query optimizations, JVM/connection-pool tuning, rate-limit or autoscaling configuration, retry/backoff changes, or escalating heavy-request patterns.
- Make the delivery path safer: Partner with squads to improve CI/CD (Jenkins/Github Actions) and GitOps deploys (ArgoCD + Helm), including health checks, readiness/liveness probes, progressive rollouts, and rollbacks for finance services.
- Inform capacity and cost: Help analyze platform usage to attribute costs to customers and workloads, inform capacity planning, and identify efficiency improvements. Partner with Cloud Operations (FinOps).
- Leave things more reliable than you found them: Improve tests, observability, and operability each time a service is touched, and grow the group's reliability guardrails into a shared, durable practice.
- General: Other duties as assigned by the Company. Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions.
What YouBring
- Experience: 3+ years in SRE, software engineering, or systems/operations roles supporting production services, with the appetite to own and grow a reliability practice.
- Systems Design: Solid understanding of systems and application design, with the ability to reason about reliability, failure modes, and performance.
- Java: Able to read, debug, and make changes to Java/Spring code today, with the willingness and aptitude to deepen JVM expertise (garbage collection, memory, thread pools) on the job.
- Cloud & Kubernetes: Experience operating services on a major cloud (ideally GCP/GKE) and exposure to Kubernetes and containers.
- Observability: Proficiency with metrics, logging, tracing/APM, and alerting.
- Database: Experience with relational database monitoring and SQL tuning (MySQL preferred), including basic indexing and query optimization.
- Automation: Proficiency in shell scripting and comfort automating routine operational tasks.
- Security & Compliance Awareness: Practical understanding of operational security for production systems, including secrets management, least-privilege access, and handling sensitive data safely in logs and telemetry.
- SLO/SLI: Familiarity with defining or operating against SLOs/SLIs, or a strong desire and aptitude to build that practice.
- Collaboration: A collaborative working style with the ability to influence and enable other engineers.
- Problem Solving: Ability to prioritize, work independently, and focus on simple, efficient, and reliable solutions.
- Education: B.S. in Computer Science or a related field, or equivalent practical experience.
Preferred Qualifications
- Spring Ecosystem: Experience with the Spring/Spring Boot ecosystem and JVM application servers.
- Batch/Workflow Orchestration: Experience with tools like Quartz, Temporal, or similar.
- Event-Driven Messaging: Familiarity with Pub/Sub or Kafka and service-to-service integration (REST/Feign, gRPC).
- Telemetry & Metrics Analysis: Experience with Prometheus/PromQL, Grafana, and log analytics.
- AI-assisted engineering: Familiarity with using AI tools to assist in engineering workflows.
- Tooling: Familiarity with secrets management (Vault, GCP Secret Manager, SOPS), code-quality gates (SonarQube), and CI/CD + GitOps (Jenkins, ArgoCD/Helm).
Benefits and Perks
At impact.com, we believe that when you’re happy and fulfilled, you do your best work. That’s why we’ve built a benefits package that supports your well-being, growth, and work-life balance.
- Flexible Working: Our Responsible PTO policy means you can take the time off you need to rest and recharge. We're committed to a positive work-life balance and provide a flexible environment that allows you to be happy and fulfilled in both your career and your personal life.
- Health and Wellness: Your well-being is a priority. Our mental health and wellness benefit includes up to 12 fully covered therapy/coaching sessions per year, with additional dependent coverage. We also offer a monthly gym reimbursement policy to support your physical health.
- A Stake in Our Growth: We offer Restricted Stock Units (RSUs) as part of our total compensation, giving you a stake in the company's growth with a 3-year vesting schedule, pending Board approval.
- Investing in Your Growth: We’re committed to your continuous learning. Take advantage of our free Coursera subscription and our PXA courses.
- Parental Support: We offer a generous parental leave policy, 26 weeks of fully paid leave for the primary caregiver and 13 weeks fully paid leave for the secondary caregiver.
- Technology Financial Support: We provide a technology stipend to help you set up your home office and a monthly allowance to cover your internet expenses
impact.com is proud to be an equal opportunity workplace. All employees and applicants for employment shall be given fair treatment and equal employment opportunity regardless of their race, ethnicity or ancestry, color or caste, religion or belief, age, sex (including gender identity, gender reassignment, sexual orientation, pregnancy/maternity), national origin, weight, neurodivergence, disability, marital and civil partnership status, caregiving status, veteran status, genetic information, political affiliation, or other prohibited non-merit factors.
_CapeTown
- ...Location: Remote Employment Type: Full-Time Industry: Cloud Infrastructure | DevOps | Site Reliability Engineering | Data Technology WatersEdge Solutions is partnering with a growing technology business to appoint a DevOps / SRE Cloud Engineer (Site Reliability...
- ...Purpose of role We’re looking for an AWS Site Reliability Engineer (SRE) to help us build and operate highly reliable, secure, and scalable cloud platforms. This role is ideal for someone who thrives at the intersection of software engineering, cloud infrastructure, and...
- ...We're seeking a talented Reliability Engineer to join our client's team in Northern Cape ! Requirements B Engineering or B Tech degree in Mechanical/Electrical Engineering Government Certificate of Competency (Mines and Works) Diploma/Degree in Project Management...
- ...A well-established mine based in Limpopo is currently seeking an experienced and qualified Reliability Engineer to join their highly engaged and dynamic workforce. Minimum Requirements Degree / Diploma in Engineering (Electrical or Mechanical) Relevant engineering...
- ...We are looking for a skilled Site Engineer to join our team in Cape Town. The Site Engineer will be responsible for overseeing and managing all construction activities on site, ensuring that the project is completed on time, within budget, and to the highest standard...
- ...client, who offers the highest quality construction, project management and specialist subcontracting services, is searching for a Site Engineer to join their team in Cape Town. ????????????????: Manage site teams by assigning specific tasks and monitoring output...
- ...collaborating hard, and holding each other to high standards while leading with empathy and kindness. The Role As a Core Reliability Engineer, you will not be joining a traditional 24/7 operations team. Instead, you will act as a central software team enabler,...
- Hire Resolves client is looking for a Site Engineer/Technician to join their team in Cape Town, for a 2 year contract starting end of August 2025 . This company is responsible for a variety of professional engineering services to provide our clients with customized solutions...
- ...Key Responsibilities Engineering & Equipment Reliability Lead all equipment maintenance and reliability initiatives across the manufacturing site. Develop and implement preventative and planned maintenance programmes. Improve equipment uptime, reliability,...
- ...Our client is one of South Africa’s leading construction companies and is looking for a driven, energetic and experienced Site Engineer to join their Building Division in Cape Town. This is an opportunity to take ownership of project execution, coordinate site activities...
- ...Site Engineer A Site Engineer role with real influence on how a busy cosmetics manufacturing site runs day to day, keeping equipment reliable, operations efficient and the working environment safe and compliant. ABOUT THE ROLE This Site Engineer position leads...
- ...Reference: RE000015-NS-1#SHIFTINTOHIGHCAREER by joining a Reputable Engineering Company in the Renewable Energy sector that seeks the expertise of an onsite Site Engineer | Cape Town Minimum Requirements Must have a minimum of 5 years experience as a Site Engineer...
- ...Job Title: Site Engineer Location: Cape Town Contract Type: Limited Duration Contract (LDC) Industry: Construction Salary :Market-related salary, based on experience and qualifications The Site Engineer will be responsible for delivering quality construction...
- ...Posted recently. We believe great workplaces are built on great people. We are looking for a Site Engineer in Cape Town, and we encourage applicants from across Western Cape to consider this opportunity. About This Position This Site Engineer role has been briefed...
- ...We are an air systems engineering company. We focus primarily on large turnkey projects, and manage the entire development process, from... ...sectors We are currently seeking a proactive and experienced Site Manager to lead one of our two site installation teams. In this...
- ...Job Title: Site Agent Hire Resolve's client is currently seeking a Site Agent to... ...Requirements: - Bachelor's degree in Civil Engineering, Construction Management, or related... ...of a team - Valid driver's license and reliable transportation Benefits: Salary:...
- Overview Ensuring that Teleperformance IT support procedures are followed. 1st line IT support, answering helpdesk calls and allocating resources. Supporting Windows 10 endpoints and Microsoft Office 365. Laptop maintenance, OS imaging & rebuilding. Deploying Anti-virus...
R 300,000 - 420,000 pa
A reliable renewable energy EPC firm is looking for Solar Site Manager to join their team in Cape Town and Benoni. Responsibilities Oversee the operations... ...Requirements Experience in construction, engineering, or energy sectors is highly desirable Strong...- ...Are you an experienced Site Manager with a strong background in electrical engineering and project supervision? A respected and forward-thinking company in the electrical contracting sector is seeking a skilled professional to lead on-site operations across a variety...
- ...stakes require permanent coordination with business entities, the Branch and Company. Activités This is a skilled, site based civil engineering position responsible for the supervision and management of civil engineering activities during construction of renewable...
- ...ATT JOB DESCRIPTION - SITE SUPERVISOR: NAME OF ROLE: Site Supervisor DEPARTMENT:Technical Department REPORTING LINE: Key Accounts Manager REPORTS:INDIRECT REPORTS: Financial Responsibilities:Col21.Weekly WagesSubmission of correct weekly wages to HR for...
- A well-established and fast-growing property development firm is looking for a highly competent Site Supervisor / Construction Manager to join their dynamic team in Cape Town. This role is ideal for someone with a strong construction background, excellent leadership skills...
- ...Strong background in site management, thorough knowledge of building processes, and a hands-on approach to managing teams and timelines... ...and leadership abilities. Bachelors degree in civil engineering or related field. Preferably based in Cape Town. If not, relocation...
R 350,000 - 500,000 pa
...A well-established construction company is currently in search of a Site Agent to join their team in Cape Town. They are in search of a candidate that specializes in High End / High Rise Residential Developments. Requirements ~5- 10 years experience ~ SACPCMP registration...R 500,000 - 700,000 pa
...that focuses on Residential, Education, Healthcare and Hospitality projects. They are urgently seeking the expertise of a Construction Site Manager in Cape Town. Responsibilities: Manage the construction site and all related aspects in Cape Town, including...- ...Site based role in Kroonstad (Free State province) As the Construction Site Manager, you will be the point of contact for all site and... ...of the site quality procedures, together with the Resident Engineer and Quality Controllers. Participating in safety and quality...
- ...takealot.com , South Africa's leading online retailer, is looking for a highly talented Senior Site Merchandiser to join our team in Cape Town. We are a young, dynamic, hyper growth company looking for smart, creative, hard-working people with integrity to join us...
- ...reputable building and civil contractor is seeking an experienced Site Agent to join their team on one of their latest projects in... ...documentation. Requirements National Diploma in Building or Civil Engineering or a Degree Construction Management or Civil Engineering....
- ...A leading construction firm is seeking an experienced Site Manager to oversee high-rise, commercial building projects. About the Role Manage day-to-day site operations, ensuring projects are delivered on time, within budget, and to the highest quality standards....
R 360,000 - 420,000 pa
...building contractor is seeking an experienced Site Manager to join its Cape Town team.... .... Valid driver's licence and own reliable transport. Proficient in Microsoft Office... ...on select market segments, namely Engineering, Finance, Supply Chain, Manufacturing, Information...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!
