Providence India
This role is sourced from an external job board and is not a ProSculpt-verified employer. Review the listing carefully, never pay any fee to apply, and verify the company before sharing personal details.
We are seeking a highly experienced and hands-on Principal Software Engineer to help define and build the next generation of our enterprise Observability and Operational Intelligence Platform . This role combines deep full-stack engineering expertise, platform architecture leadership, product strategy, and modern observability practices. The ideal candidate will have experience owning product direction, platform architecture, and technology strategy for enterprise-scale engineering solutions. Prior experience building observability, monitoring, operational intelligence, AIOps, or telemetry-driven platforms is highly desirable. Own and drive the technical product vision, platform architecture, and technology strategy for a cloud-native observability and intelligence platform. Lead architecture, design, and development of scalable, resilient, and highly available enterprise solutions. Build modern user experiences using React, TypeScript, and contemporary front-end frameworks . Design and develop backend services, APIs, microservices, telemetry pipelines, and intelligence services. Define platform standards, technology selections, and engineering best practices aligned with long-term product strategy. Design and optimize large-scale telemetry ingestion, aggregation, correlation, and analytics pipelines supporting real-time operational intelligence. Drive innovation across Monitoring, Observability, Event Correlation, AIOps, Operational Analytics, Predictive Intelligence, and Self-Healing Automation . Evaluate emerging technologies, market trends, and customer needs to continuously evolve product capabilities and roadmap direction. Partner with engineering, operations, and business stakeholders to translate enterprise challenges into scalable platform solutions. Establish engineering excellence across reliability, performance, automation, security, scalability, and developer experience. 10+ years of software engineering experience with strong expertise in full-stack product development. Proven experience owning and delivering technical product direction, platform architecture, and enterprise technology strategy.Expert-level proficiency in React, TypeScript, JavaScript, HTML/CSS , and modern UI architectures.Strong backend engineering experience using .NET, Java, Node.js, or Python .Demonstrated experience designing enterprise-scale architectures, distributed systems, microservices, event-driven architectures, and cloud-native platforms.Strong experience with Azure (preferred), AWS, or GCP , Kubernetes, CI/CD, DevOps, Infrastructure as Code, and platform engineering practices.Experience designing and optimizing large-scale telemetry ingestion, processing, event correlation, and analytics architectures supporting real-time operational intelligence.Deep understanding of observability concepts including Metrics, Logs, Traces, Dashboards, Alerting, SLO/SLI Management, Incident Management, Operational Analytics, and Platform Reliability Engineering .Experience with leading observability platforms such as Datadog, Dynatrace, Grafana, Splunk, New Relic, AppDynamics, Elastic, OpenTelemetry, or similar technologies .Experience designing AI-powered intelligence layers leveraging:Machine Learning and Predictive AnalyticsAgentic AI WorkflowsLLM-based Operational AssistantsRoot Cause Analysis and Remediation Recommendations Experience translating customer feedback, operational insights, and industry trends into platform capabilities and product roadmap investments. Strong product mindset with a focus on customer value, business outcomes, innovation, and engineering excellence. Preferred Experience Experience building enterprise Observability, AIOps, Operational Intelligence, Reliability Engineering, IT Operations, Monitoring, or Platform Engineering products . Experience designing systems that process and analyze high-volume telemetry data at scale. Experience building intelligent platforms that combine telemetry, AI/ML, automation, and operational workflows. Exposure to SRE practices, self-healing automation, anomaly detection, capacity prediction, and operational decision intelligence. .
Want a personalised feed?
Sign up to save jobs, track applications, and get AI-matched recommendations.
Create a free account