Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.If you'd like to build the world's best AI cloud, join us.We are seeking a Senior Incident Manager to lead critical incident response across our AI data center infrastructure. This role is responsible for coordinating rapid resolution of service-impacting events, improving operational resilience, and driving incident management best practices across infrastructure, networking, platform engineering, and data center operations.Role OverviewThe Senior Incident Manager is responsible for leading the end-to-end lifecycle of operational incidents impacting AI infrastructure and data center services. This individual acts as the central command point during major incidents, ensuring rapid triage, cross-team coordination, effective communication, and structured post-incident analysis.This role requires deep operational expertise in high-availability infrastructure, large-scale GPU clusters, networking, and cloud platforms, along with strong leadership and communication skills.What You’ll DoIncident LeadershipLead the response to critical (SEV-1 / SEV-2) incidents impacting AI infrastructure, GPU clusters, networking, storage, and data center operations.Serve as the Incident Commander during major outages, coordinating engineering, networking, facilities, and vendor teams.Act as the liaison between leadership and external teams during incidents / post-incidents to provide updates and status summaries.Establish clear incident timelines, triage actions, and resolution plans.Incident Management OperationsOwn the incident response lifecycle including:Assisting Technical TriageEscalationCoordinationResolutionPost-incident reviewEnsure timely and accurate communication with internal stakeholders and leadership.Maintain incident response documentation and operational playbooks.Conduct analysis on incidents and identify patterns / trends for improvement in response and systems reliability.Work in an On-Call Rotation to respond to, lead, and coordinate incidentsCross-Functional CoordinationWork closely with:Data center operationsInfrastructure engineering & operationsNetwork engineeringPlatform reliability engineeringSecurity operationsHardware and facility vendorsDrive alignment during outages involving multiple infrastructure layers.Post-Incident Analysis & Continuous ImprovementLead post-incident reviews (PIRs) and root cause analysis. Identify systemic reliability gaps and implement corrective actions.Track incident metrics including MTTR, MTTD, and incident recurrence rates.Operational ExcellenceImprove incident response processes, escalation paths, and tooling by working with technical support and engineering teams..Contribute to runbooks, operational standards, and reliability frameworks.Support implementation of automation and observability improvements.Communication & ReportingProvide executive-level incident summaries and reports.Deliver clear, concise updates during active incidents.Maintain incident dashboards and operational health reporting.You8+ years experience in incident management, site reliability engineering, or infrastructure operationsExperience managing incidents in large-scale distributed infrastructure environmentsStrong understanding of:Data center operationsGPU compute clustersNetworking and storage infrastructureCloud or hybrid infrastructure platformsProven ability to lead high-pressure incident response situationsExperience with incident management frameworks (ITIL, SRE, or equivalent)Excellent communication and stakeholder management skillsExperience with incident tracking and monitoring tools such as:PagerDutyServiceNowJiraDatadogPrometheus / GrafanaNice to HaveExperience operating AI or HPC infrastructureBackground in SRE, infrastructure engineering, or data center operationsFamiliarity with high-density GPU environments (NVIDIA clusters, InfiniBand networks)Experience with hyperscale or colocation data center environmentsKnowledge of automation and incident response toolingKnowledge of and experience with Incident command system (ICS)Experience in leading and developing incident command from stractchKey CompetenciesIncident Command & LeadershipOperational Decision MakingCross-Team CoordinationRoot Cause AnalysisCrisis CommunicationInfrastructure ReliabilityWhat Success Looks Like in This RoleReduced Mean Time to Resolution (MTTR) for critical incidentsImproved cross-team incident coordinationHigh-quality post-incident reviews and corrective actionsIncreased infrastructure reliability and operational maturitySalary Range InformationThe annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.About LambdaFounded in 2012, with 500+ employees, and growing fastOur investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent CoveWe have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOGOur values are publicly available: We offer generous cash & equity compensationHealth, dental, and vision coverage for you and your dependentsWellness and commuter stipends for select roles401k Plan with 2% company match (USA employees)Flexible paid time off plan that we all actually useEqual Opportunity EmployerLambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.Compensation Range: $125K - $195KLocationRemote, USA; San Jose Office (Zanker)Employment TypeFull timeLocation TypeRemoteDepartmentData Center BusinessCompensationRemote, USA$125K – $166KSan Jose$146K – $195K
...This is a hands-on leadership role: you will set the direction and also roll up your sleeves. Youll own the company's multi-year IT roadmap, budget, policies, and governance frameworks -- serving as the senior technology advisor to executive leadership and translating...
...facility near you! With one centralized pickup location and smaller delivery zones, Gopuff makes earning effortless. It's simple: deliver... ...medications to food, drinks and more. Sign up to be a Gopuff Bike Delivery Partner today and experience the easiest way to earn big...
...Cross Country Education is seeking a local contract School Physical Therapist for a local contract job in La Plata, Maryland. Job Description... ...Maryland State Board of Physical Therapy Experience with school-based evaluations, IEP/IFSP development, multidisciplinary...
Columbia Mental Health seeks a Behavioral Health Advanced Practice Provider to deliver comprehensive psychiatric evaluations, medication management, and collaborative care in Washington, DC. Youll treat diverse adults with mood, anxiety, and other mental health conditions...
Music & Arts seeks a passionate Music Teacher to provide highquality private and group lessons in our retail store. In this student... ..., educationdriven culture that values community engagement and offers opportunities to grow in teaching and the music retail industry.