Principal Production Engineer — Operations
- Lilly
- Telangana, India
- INR 3,000,000 – INR 5,000,000
At Lilly, the work is demanding because patients are waiting. We unite caring with discovery to help make life better for people around the world, knowing that every decision, every detail, and every day matters. Headquartered in Indianapolis, Indiana, our over 50,000 employees around the globe take on complex challenges to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. This is hard, urgent, selfless work—but it’s work worth doing. If you’re driven by purpose and ready to bring your best to work that truly matters for patients, we invite you to join us.
Principal Production Engineer — Operations
Job description
About Lilly
At Lilly, everything we do starts with patients. We unite caring with discovery to make life better for people around the world. Headquartered in Indianapolis, Indiana, our global team of over 50,000 employees work with urgency and purpose to discover and deliver life-changing medicines, strengthen how health is understood and managed, and support the communities we serve. We bring our best to this work because people depend on it. If you’re driven by purpose and determined to make a meaningful difference for patients, we invite you to bring your skill and your commitment to Lilly.
About Tech@Lilly
At Lilly, technology is not a support function. It is how a global medicine company operates, innovates, and delivers. Lilly in Hyderabad builds the capabilities that make this possible, cloud platforms, AI systems, and automation at enterprise scale, all in service of a purpose that makes this technology work genuinely distinctive, from advancing drug discovery to enabling connected clinical trials to keeping a global medicine company running at the standard patients deserve.
Experience
10+ years
Location
Hyderabad (Onsite)
Employment Type
Full-time
Job Family
R3 — Operations / Application Support / Production Engineering / SRE-aligned Support
Time Allocation
Majority operational coverage and incident command · 20–30% engineering and toil reduction
About the technology organization
Technology at Lilly builds and operates mission-critical digital products and platforms that support the discovery, development, and delivery of medicines that make life better for people around the world. Our teams operate in highly regulated, high-availability environments, where operational excellence, reliability, and quality are non-negotiable.
Within Tech@Lilly, the Digital Core organization applies a product, platform, and reliability-first mindset, ensuring that operational capabilities scale sustainably across the enterprise.
About the Team
Tech@Lilly builds and maintains capabilities using pioneering technologies like the most prominent tech companies. What differentiates Lilly IT is that we redefine what's possible through tech to advance our purpose, creating medicines that make life better for people around the world, including data-driven drug discovery, connected clinical trials, resilient enterprise platforms, and intelligent digital operations. We hire the best technology professionals from a variety of backgrounds, so they can bring an assortment of knowledge, skills, and diverse thinking to deliver creative solutions in every area of our business.
The Digital Core team leads Lilly's transformation into the Digital and AI era. They inspire digitally empowered teams to new ways of working and accelerate innovation and agility. This team powers and advances the entire company by building and maintaining world-class technology capabilities and platforms.
Lilly Capability Centre India (LCCI), Hyderabad, is Lilly's premier Global Technology Hub, harnessing data, AI, analytics, and digital solutions to revolutionize healthcare and improve patient outcomes worldwide.
The Production Reliability Engineering team is the engineering-first function that owns operations, incident command, runbook automation, and the safe evaluation of agent-assisted remediation across a multi-application production estate. The estate is deliberately diverse: it spans low-code platforms (Power Platform, Power Automate), SaaS ecosystems (Salesforce / SFDC, SharePoint), traditional enterprise stacks (Java/JVM, .NET, Node.js, Python), automation platforms (RPA — Automation Anywhere, UiPath-class tooling), data and analytics platforms (Power BI, Tableau, relational and NoSQL databases), and the integration layer that ties them together. The team works in close partnership with the engineering team that builds the agentic automation platform that increasingly reduces manual operational load.
Role summary
As a Principal Production Engineer in Operations, you are the final-escalation authority within your shift and the named on-call anchor for the team. You drive root-cause analyses to engineering-grade quality before handoff to the SRE engineering practice in the same pillar, and you are the last call before the engineering lead is engaged.
What sets this seat apart from a single-stack senior engineer is the breadth of technology you operate across. The estate covers low-code platforms, SaaS ecosystems, traditional enterprise stacks, RPA and automation tooling, integration middleware, and data and analytics platforms. You are expected to debug across all of them — not by memorising every product, but by reading logs, traces, code, and configuration in any reasonable stack and isolating where it broke. Candidates who thrive in this role pick up unfamiliar tech in weeks, not quarters.
This is a senior individual-contributor role with real shift weight and real authority. The team supports continuous operations across multiple shifts, and you will anchor on-call coverage for your assigned shift while partnering with the engineering lead on cross-shift incident command.
What you'll be doing
- Shift anchor and on-call leadership
- Hold the on-call anchor role for your shift and lead the L1/L2/L3 operations engineers working alongside you; carry the page that the engineering lead does not.
- Direct triage and dispatch in real time: pull L1 and L2 to the work they can own, keep senior capacity on the hard problems, and step in only where the call genuinely needs you.
- Make incident-disposition decisions within your shift (escalate, resolve, or open a problem record) with disciplined judgment, and coach your team to make the same calls when you are not on the bridge.
- Run effective shift handoffs to maintain operational continuity across globally distributed teams, and raise the bar on handoff quality so the next shift inherits clarity, not noise.
- Multi-stack debugging and major incident response
- Be the final-escalation point before the engineering lead, and earn the right to be the last call by consistently isolating fault domains across the stack faster than anyone else on shift.
- Debug across very different technology families — low-code (Power Platform, Power Automate), SaaS (Salesforce/SFDC, SharePoint), enterprise stacks (Java/JVM, .NET, Node.js, Python), RPA (Automation Anywhere or equivalent), integration middleware, and data/analytics platforms (Power BI, Tableau, SQL and NoSQL stores) — using logs, traces, metrics, code reading, and first-principles reasoning. Deep domain knowledge of every product is not the bar; the bar is finding where it broke and stopping the bleed.
- Lead in-shift major incident execution and partner with the engineering lead on cross-shift incident command. Own the bridge, drive comms, and keep the response moving.
- Coordinate cross-team response across engineering, product, platform, security, and vendor teams when needed, and hold each owner to clear next steps and timestamps.
- Root-cause engineering and pattern feedback
- Drive root-cause analyses to engineering-grade quality before handoff to the SRE engineering practice in the same pillar, covering symptom, timeline, fault domain, root cause, contributing factors, and the engineering work that follows.
- Partner with the engineering lead on problem management for chronic incident patterns, and convert recurring incidents into runbooks, automation, or alert tuning that L1 and L2 can execute on their own next time.
- Identify cross-product, cross-stack failure patterns across application, integration, infrastructure, data, and identity layers, and feed them upstream to the right engineering team with enough evidence that the fix is unambiguous.
- Partner with the agentic automation engineering team on which production patterns become safe agent-assisted remediations — including the guardrails, confidence thresholds, and human-in-the-loop checkpoints that keep automation honest.
- Operational readiness, runbook engineering, and toil reduction
- Validate work packages from the agentic automation engineering team before they go live in production. Pressure-test runbooks, alerting, rollback paths, and observability coverage.
- Write and maintain runbooks, scripts, and ServiceNow workflows that let L1 and L2 resolve routine incidents without paging up. Measure your impact in the toil you remove from the shift, not just the incidents you personally close.
- Maintain ServiceNow lifecycle hygiene across the supported application estate.
- Provide senior operational oversight for releases and deployments within your shift, including go/no-go calls when readiness gaps surface late.
- Compliance, security, and regulated-environment readiness
- Ensure operational practices comply with Lilly standards and applicable regulatory requirements, including change control, access reviews, audit trail, and evidence capture during incidents.
- Promote secure handling of sensitive data during incidents, including the data that surfaces in logs, screenshots, and bridge transcripts.
How you will succeed
At the principal level in operations, success is defined by named-owner posture and sustained operational outcomes:
- You judge escalation timing well; too early erodes credibility, too late costs the business.
- You write root-cause analyses that leadership reads.
- You hold the operational bar consistently across your shift.
- You move comfortably across very different technology families, and L1/L2 engineers learn how to debug across stacks from working with you.
- Toil you remove from the shift shows up in fewer pages, faster handoffs, and L1/L2 closing incidents they would have escalated last quarter.
What you should bring
Required
- 10+ years of professional technology experience, with at least four years in production incident response and named on-call ownership.
- Demonstrated ability to debug across multiple technology stacks without deep application-domain knowledge. You read logs, traces, code, and configuration across at least four of the following families and have actually shipped fixes in three: low-code platforms (Power Platform, Power Automate, or equivalent); SaaS ecosystems (Salesforce/SFDC, SharePoint, ServiceNow customisations); traditional enterprise stacks (Java/JVM, .NET, Python, Node.js); Linux and Windows server; relational and NoSQL databases; REST and messaging integration layers; RPA platforms (Automation Anywhere, UiPath, or Blue Prism); data and analytics platforms (Power BI, Tableau, or equivalent).
- Working fluency with at least one major cloud platform (AWS preferred, Azure or GCP acceptable), including networking, IAM, compute, storage, and the operational footprint of managed services.
- Working fluency with at least one major observability platform (Splunk preferred, Datadog, Dynatrace, New Relic, Elastic, or equivalent). You build dashboards and queries during incidents, not just read them.
- Scripting (Python or Bash, ideally both) good enough to write quick remediations and one-off automations during an incident.
- Proven experience leading other operations engineers, including running shifts, mentoring L1 and L2, and lifting team capability over time.
- Strong ServiceNow operational fluency and deep ITIL practice across incident, problem, change, and CMDB hygiene.
- Visible willingness to learn unfamiliar stacks fast — the candidates who thrive in this role pick up new tech in weeks, not quarters.
- Bachelor's degree or higher in Computer Science, Information Technology, or a closely related field.
Preferred
- Experience operating in highly regulated industries (pharma, financial services, healthcare, or similar) where audit, validation, and change control are non-negotiable.
- Breadth across additional stack layers such as container platforms (Kubernetes, OpenShift, ECS), infrastructure-as-code (Terraform, Ansible), CI/CD pipelines, identity providers, API gateways, and integration middleware (MuleSoft, Boomi, Workato, or equivalent).
- Exposure to AIOps tooling and agent-assisted operational workflows, and a point of view on where automation pays off versus where it adds risk.
- SLI/SLO/error-budget practice; can define meaningful indicators for an unfamiliar service and defend the targets.
- Track record of root-cause analyses you can describe end to end, covering symptom, timeline, fault isolation, root cause, contributing factors, and the engineering work that followed.
- Formal certifications welcome but not required, such as ITIL, AWS, Azure, CKA, or SRE Foundation.
Leadership expectations
- Brings the named-owner posture this role demands.
- Drives accountability, clarity, and calm during high-pressure situations.
- Treats ServiceNow as the operational system of record, not as an annoyance.
- Raises the capability of the team through example, especially in how to debug across unfamiliar stacks.
- Treats picking up a new technology family as part of the job, not as a special project.
Additional information
Availability to work flexible work hours is/may be required. This team supports continuous operations and may require non-standard work hours, including some work on weekends and holidays. Appropriate adjustments in benefits will be provided for employees working non-standard hours where applicable.
This is an onsite role based in Hyderabad. Candidates should be open to working different shifts (e.g., 6 AM to 2 PM and 2 PM to 11 PM) when required to align with global delivery partners.
At Lilly, caring is not only what we do for patients. It is how we work. We believe the people who dedicate themselves to making medicines better deserve an environment that makes their lives better too, one where they are supported, respected, and given the space to do their best work. This is not just a policy. It is who we are.
Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form for further assistance.
Lilly does not discriminate on the basis of age, race, color, religion, gender, sexual orientation, gender identity, gender expression, national origin, protected veteran status, disability, or any other legally protected status.
Lilly is dedicated to helping individuals with disabilities to actively engage in the workforce, ensuring equal opportunities when vying for positions. If you require accommodation to submit a resume for a position at Lilly, please complete the accommodation request form (https://careers.lilly.com/us/en/workplace-accommodation) for further assistance. Please note this is for individuals to request an accommodation as part of the application process and any other correspondence will not receive a response.
Lilly does not discriminate on the basis of age, race, color, religion, gender, sexual orientation, gender identity, gender expression, national origin, protected veteran status, disability or any other legally protected status.
#WeAreLilly
Skills
- Process Engineering
- Manufacturing Operations
- GMP Compliance
- Root Cause Analysis
- Project Management
- continuous improvement
- Cross-functional Collaboration








