Senior L1
Save this job and keep your search organized
Create a free account to save jobs, create alerts and return to this listing from your dashboard.
Key Responsibilities
Production Support & Incident Management
. Act as the primary Ll /L2 support contact for digital platforms, e-commerce systems, websites, and customer-facing services.
. Monitor incident queues, service requests, alerts, and support tickets, ensuring adherence to SLAs and operational procedures.
. Lead incident triage, troubleshooting, escalation, and resolution activities.
. Perform impact assessment and coordinate with relevant stakeholders during service disruptions.
. Support incident management activities and facilitate communication during critical outages.
. Conduct post-incident reviews and root cause analysis (RCA) to prevent recurrence.
. Develop and maintain operational runbooks, support procedures, and knowledge base articles.
System Monitoring & Reliability
. Monitor application, infrastructure, and business service health using observability and monitoring tools.
. Analyse system performance, availability, error trends, and capacity utilization.
. Configure and tune alerts to reduce noise and improve operational visibility.
. Collaborate with engineering teams to improve system reliability and operational resilience.
. Support Site Reliability Engineering (SRE)practices, including reliability metrics, incident reduction, and service availability improvements.
Cloud & Infrastructure Support
. Provide operational support for cloud-hosted applications and infrastructure, primarily on AWS.
. Perform first-level troubleshooting on:
o Compute services (EC2, ECS, Lambda)
o Networking
o Load Balancers
o CDN services
o Storage services
. Investigate infrastructure-related issues affecting application performance or availability.
. Support deployment verification and post-release monitoring activities.
Application & Integration Support
. Troubleshoot application issues across web, mobile, APIs, and middleware platforms.
. Analyse application logs, monitoring data, and system traces to identify root causes.
. Support integrations with external systems, partners,