At Early Warning, we’ve powered and protected the U.S. financial system for over thirty years with cutting-edge solutions like Zelle®, Paze℠, and so much more. As a trusted name in payments, we partner with thousands of institutions to increase access to financial services and protect transactions for hundreds of millions of consumers and small businesses.
Positions located in Scottsdale, San Francisco, Chicago, or New York follow a hybrid work model to allow for a more collaborative working environment.
Candidates responding to this posting must independently possess the eligibility to work in the United States, for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.
Staff Software Engineer - AIOps
Overall Purpose
The Staff Software Engineer - AIOps is a senior hands-on technical contributor responsible for designing, building, and operating the AIOps agent platform that enables AI agents to observe events across Early Warning technology environments, generate diagnoses and recommendations, and execute approved actions within defined safety controls. The role focuses on the software systems, control loop, tool interfaces, telemetry, evaluation capabilities, and operator experience required to run agentic capabilities safely and reliably in production; it is not primarily a model development or research role.
This position executes complex technical work with general direction, independently makes implementation decisions within established architecture and standards, and may seek guidance for novel, highly ambiguous, or cross-enterprise decisions. The Staff Engineer applies software engineering, distributed systems, and platform engineering principles and partners across Engineering, Technology Operations, Security, Risk and Compliance, and other technology teams to establish reusable patterns and guardrails for progressive levels of automation.
Essential Functions
- Apply mature software engineering practices across the agent runtime, tool layer, APIs, and operator experience, including versioned interfaces, automated testing, code review, release management, observability, and regression coverage.
- Execute complex AIOps engineering assignments with general direction; break work into deliverable components, identify practical implementation solutions, raise risks and dependencies, and seek guidance when decisions extend beyond established architecture or standards.
- Contribute to the design and evolution of the agent platform architecture; maintain architecture decision records, define interface standards, and evaluate significant design and build-versus-buy tradeoffs with appropriate technical guidance.
- Design, build, test, and operate the core agent runtime, including event ingestion, context assembly, planning, tool invocation, verification, escalation, and response handling.
- Design and maintain scoped, typed, and auditable tool interfaces that enable agents to interact with CI/CD platforms, infrastructure automation, Kubernetes, network services, ITSM platforms, observability systems, and secrets management solutions.
- Define and implement controls for agent actions, including read-only and advisory capabilities, human approval requirements, narrowly scoped unattended actions, permission boundaries, blast-radius limits, dry-run capabilities, reversibility, and emergency shutdown controls.
- Define tool contracts, versioning standards, testing requirements, and intended behavior so agent-accessible capabilities are managed as reliable production interfaces.
- Build event-routing capabilities that receive and prioritize alerts, pipeline failures, tickets, operational requests, and other technology events and provide appropriate context to the agent runtime.
- Establish telemetry capabilities used by the agent, including logs, metrics, and traces from relevant technology platforms, and contribute to standards for telemetry quality and schema evolution.
- Design and operate model routing, session and state management, retries, timeouts, and cost and latency controls appropriate for a production service.
- Implement comprehensive auditability for events received, decisions generated, approvals obtained, and actions executed to support operational, risk, compliance, and examination requirements.
- Develop evaluation capabilities for agent decision quality, including historical incident replay, controlled or shadow-mode evaluation, measurement of proposed actions, and regression testing.
- Build feedback mechanisms that incorporate human approvals, overrides, and operational outcomes into measurable improvements to platform quality.
- Package reusable agent capabilities, tool integrations, APIs, documentation, and implementation patterns so other technology teams can extend and adopt the platform through defined self-service practices.
- Develop and maintain the operator experience, including approval workflows, agent activity and decision history, telemetry views, and APIs that expose agent state to user interfaces and other systems.
- Partner with Engineering, Technology Operations, Architecture, Security, Risk, and Compliance stakeholders to evaluate technical tradeoffs and ensure solutions meet reliability, security, operational, and regulatory requirements.
- Define and monitor measures of platform effectiveness, including recommendation quality, approval and override rates, action success, rollback frequency, latency, cost, adoption, and operational outcomes.
- Support the company's commitment to risk management and protecting the integrity and confidentiality of systems and data.
Minimum Qualifications
- Education and/or experience typically obtained through a bachelor's degree in computer science, engineering, or a related technical field.
- Typically eight or more years of related experience in software engineering, platform engineering, distributed systems, DevOps, AIOps, or a related technical discipline.
- Demonstrated strength in software engineering and distributed-systems design, including the ability to turn ambiguous technical problems into maintainable, testable solutions and reason about interfaces, state, retries, timeouts, idempotency, failure modes, and operational risk.
- Demonstrated ability to execute complex technical work with general direction, independently make implementation decisions within established architecture and standards, and recognize when guidance or escalation is needed for novel or enterprise-wide decisions.
- Demonstrated customer-first approach, balancing internal customer and operator needs with safety, security, reliability, explainability, and control requirements.
- Ability to respectfully challenge assumptions with data and technical reasoning, then disagree and commit by supporting and executing the final decision.
- Strong production development experience with Python, including services, APIs, automation, orchestration systems, or agent runtimes.
- Hands-on experience developing production agentic systems, tool-using AI applications, or autonomous or semi-autonomous software control loops beyond proof-of-concept implementations.
- Experience designing tool-use or function-calling interfaces for AI agents using MCP or comparable integration patterns.
- Experience implementing production controls for AI or automated systems, including evaluation, guardrails, monitoring, cost and latency management, and management of unintended or failed actions.
- Experience with event-driven architectures and messaging technologies such as Kafka, Pub/Sub, RabbitMQ, or comparable platforms.
- Experience with production observability practices, including structured logging, metrics, distributed tracing, and technologies such as OpenTelemetry, Prometheus, or Grafana.
- Experience developing front-end solutions using modern frameworks such as React, Next.js, or Angular, along with backend or API development using TypeScript or Python.
- Demonstrated ability to work effectively in an environment requiring strong controls, explainability, approval processes, and auditability for automated actions.
- Demonstrated ability to lead complex technical work, collaborate across engineering and operations disciplines, and influence implementation decisions without direct authority.
- Strong written and verbal communication skills, including technical documentation, architecture decision records, implementation standards, and control documentation.
- Background and drug screen.
Preferred Qualifications
- Experience with agent orchestration frameworks such as LangGraph or experience developing comparable orchestration capabilities directly.
- Go development experience.
- Experience with CI/CD platforms, infrastructure automation, Kubernetes operations, ITSM platforms, or technology operations.
- Experience leading a complex AIOps or agent-platform capability from design through production implementation with periodic architectural or stakeholder guidance.
- Experience with anomaly detection, time-series analysis, or related machine-learning applications.
- Experience with chaos engineering, simulation, historical incident replay, or game-day testing for distributed or automated systems.
- Experience mapping automated-system controls to regulatory or control frameworks such as PCI DSS, SOX ITGC, NYDFS Part 500, or FFIEC guidance.
- Experience in FinTech, banking, payments, or another highly regulated industry.
- Current AWS, Kubernetes, AI engineering, security, or other relevant technical certification.
Physical Requirements
Working conditions consist of a normal office environment. Work is primarily sedentary and requires extensive use of a computer and involves sitting for periods of approximately four hours. Work may require occasional standing, walking, kneeling and reaching. Must be able to lift 10 pounds occasionally and/or negligible amount of force frequently. Requires visual acuity and dexterity to view, prepare, and manipulate documents and office equipment including personal computers. Requires the ability to communicate with internal and/or external customers.
Employee must be able to perform essential functions and physical requirements of position with or without reasonable accommodation.
Candidates responding to this posting must independently possess the eligibility to work in the United States at the date of hire.
The above job description is not intended to be an all-inclusive list of duties and standards of the position.
The base pay scale for this position in:
Phoenix, AZ/ Chicago, IL in USD per year is: $124,000 - $165,000.
Additionally, candidates are eligible for a discretionary incentive plan and benefits.
This pay scale is subject to change and is not necessarily reflective of actual compensation that may be earned, nor a promise of any specific pay for any specific candidate, which is always dependent on legitimate factors considered at the time of job offer. Early Warning Services takes into consideration a variety of factors when determining a competitive salary offer, including, but not limited to, the job scope, market rates and geographic location of a position, candidate’s education, experience, training, and specialized skills or certification(s) in relation to the job requirements and compared with internal equity (peers). The business actively supports and reviews wage equity to ensure that pay decisions are not based on gender, race, national origin, or any other protected classes.
- Additional Job Description
Additional Job Description
Some of the Ways We Prioritize Your Health and Happiness
Healthcare Coverage – Competitive medical (PPO/HDHP), dental, and vision plans as well as company contributions to your Health Savings Account (HSA) or pre-tax savings through flexible spending accounts (FSA) for commuting, health & dependent care expenses.
401(k) Retirement Plan – Featuring a 100% Company Safe Harbor Match on your first 6% deferral immediately upon eligibility.
Paid Time Off – Flexible Time Off for Exempt (salaried) employees, as well as generous PTO for Non-Exempt (hourly) employees, plus 11 paid company holidays and a paid volunteer day.
12 weeks of Paid Parental Leave
Maven Family Planning – provides support through your Parenting journey including egg freezing, fertility, adoption, surrogacy, pregnancy, postpartum, early pediatrics, and returning to work.
And SO much more! We continue to enhance our program, so be sure to check our Benefits page here for the latest. Our team can share more during the interview process!
Early Warning Services, LLC (“Early Warning”) considers for employment, hires, retains and promotes qualified candidates on the basis of ability, potential, and valid qualifications without regard to race, religious creed, religion, color, sex, sexual orientation, genetic information, gender, gender identity, gender expression, age, national origin, ancestry, citizenship, protected veteran or disability status or any factor prohibited by law, and as such affirms in policy and practice to support and promote equal employment opportunity and affirmative action, in accordance with all applicable federal, state, and municipal laws. The company also prohibits discrimination on other bases such as medical condition, marital status or any other factor that is irrelevant to the performance of our employees.
Early Warning Services LLC is a proud participant in E-Verify, a federal program to help ensure a legal and authorized workforce. As part of our hiring process, we electronically verify the employment eligibility of all new hires through E-Verify. For more information on your rights and responsibilities under E-Verify please visit Home | E-Verify.