The role
From Oracle's own posting.
Oracle Cloud Infrastructure (OCI) delivers mission-critical cloud services to enterprises worldwide. The Physical Networking Automation and Tooling team builds innovative solutions that improve the productivity and efficiency of Network Engineering teams through automation, observability, and actionable insights at hyperscale.
As a Principal Network Developer, you will design, build, and deliver scalable automation frameworks and advanced platforms that leverage AI/ML to drive operational excellence across OCI’s global network. You will develop event-driven data pipelines that process network signals such as failures and power capabilities including anomaly detection, automated remediation, self-healing operations, model training, and real-time inference.
Build the intelligent automation that powers OCI’s global network at hyperscale. Design AI/ML driven platforms that detect failures, automate remediation, and enable self-healing operations. Solve complex networking challenges across one of the world’s largest cloud infrastructures.
You are passionate about building software that solves real-world operational challenges. You thrive in a fast-paced environment, are comfortable working with complex distributed systems, and value simplicity, scalability, and collaboration.
Preferred Qualifications
At least 8–10 years of experience in software engineering, automation development, process-as-code systems, or a related field
Bachelor’s degree in Computer Science, Computer Engineering, a related technical field, or equivalent practical experience
Strong programming skills in Java and Python
Experience designing and building scalable distributed systems, microservices, and cloud-native applications
Experience defining technical architecture and leading complex, cross-functional engineering initiatives
Proficiency with Linux environments, scripting, and software development tools
Understanding of network operations or large-scale IT infrastructure
Strong problem-solving, organizational, and communication skills
Experience applying AI/ML to anomaly detection, operational automation, or large-scale infrastructure
Experience with model training, production inference, and event-driven data pipelines
Experience using AI-assisted development tools for code generation, testing, review, and debugging
Key Responsibilities
Network Design, Development, and Validation:
-Designs and develops advanced network systems that support scalable enterprise, data center, or cloud environments, ensuring alignment with existing infrastructure patterns.
-Validates production networks and develops scaling strategies for diverse and evolving use cases across systems or services.
-Participates in solution architecture discussions, providing technical guidance on network requirements, tradeoffs, and design decisions.
-Identifies and assesses complex risks in network design and recommends mitigation strategies before deployment.
-Collaborates with vendors and internal stakeholders to align on hardware, firmware, software, and cloud network code requirements.
-Designs playbooks for resolving common and uncommon network issues involving moderately complex systems.
-Analyzes network workflows to identify inefficiencies and proposes scalable solutions to improve performance.
Automation and Scripting:
-Designs and builds advanced modules within automation frameworks to support testing, operations, and service reliability across multiple environments.
-Automates high-impact network tasks for production and lab environments, and optimizes repetitive, cross-team workflows to improve operational efficiency.
-Develops and maintains advanced dashboards, telemetry tools, and alerting systems to enable proactive monitoring and faster issue resolution.
-Writes, enhances, and documents reusable scripts that streamline routine network operations across teams and product domains.
-Modifies and extends infrastructure pipelines and configuration management tools to meet evolving requirements using existing playbooks and custom logic.
-Independently configures and extends tools to support new products or services, ensuring alignment with platform standards and integration best practices.
Testing and Quality Assurance:
-Designs and develops advanced test strategies and reusable test cases that validate complex network systems and enhance long-term network integrity and reliability.
-Leads team-level test practices and cross-functional collaboration, helping define robust testing protocols and ensuring consistency in execution across environments.
-Implements and refines post-incident validation methods, integrating break-fix outcomes and lessons learned into continuous test improvements.
-Reviews and approves L1 and L2 network changes, and presents high-impact changes to cross-team change management boards, ensuring risks are understood and mitigated.
-Partners with cross-functional teams to drive pre-production validation efforts, ensuring systems and environments meet compliance and deployment-readiness criteria.
-Prepares and delivers audit documentation and test evidence to support internal Governance, Risk, and Compliance (GRC) processes, maintaining alignment with security and policy requirements.
Monitoring and Reliability:
-Works with others (e.g. monitoring teams) to customize dashboards, telemetry pipelines, and alerting systems, defining service level objectives (SLOs) based thresholds and metrics to monitor network health.
-Partners with monitoring and Site Reliability Engineering (SRE) teams to refine alerting systems and enhance early detection of network anomalies and degradation.
-Leads incident response during support rotations, driving root cause analysis and coordinating cross-team resolution for escalated issues.
-Builds and improves internal tools that enable frontline support teams to efficiently respond to network failures and recurring operational issues.
Cross-Team Collaboration and Leadership:
-Drives test and deployment milestones in collaboration with project/program managers, adapting plans to address risks and ensure cross-team alignment.
-Mentors junior engineers and acts as a technical SME, providing hands-on guidance in resolving complex issues.
-Leads customer engagements on technical escalations and delivers clear, actionable root cause analysis (RCA) documentation.
-Influences team-level roadmap and architecture decisions, contributing solutions that align with broader engineering goals.
-Coordinates with vendors and internal teams to resolve standards misalignment, ensuring integration meets technical and operational expectations.
Performance and Capacity Management:
-Leads analysis of network performance metrics (e.g., latency, throughput, packet loss) to identify systemic inefficiencies and drive scalable improvements.
-Forecasts infrastructure needs using performance and capacity data, ensuring systems are prepared for anticipated traffic and service growth.
Core Responsibilities
Planning & Execution:
-Manages and coordinates moderately complex tasks, monitoring timelines and deliverables to ensure timely completion and adherence to requirements for a moderately-sized project or initiative. Efficiently delegates, monitors, and prioritizes work across multiple projects, providing technical oversight and adjusting plans to address shifts in resources or timelines.
Collaboration & Partnership:
-Collaborates across the organization to align on expectations and achieve shared objectives. Leverages understanding of business leaders, stakeholders, and/or customers to ensure proposed solutions meet their needs. Supports inclusivity by actively seeking and listening to diverse perspectives, ensuring others feel heard and respected.
Problem Solving:
-Identifies and addresses moderately complex issues by analyzing a wide range of data and/or information to identify solutions in accordance with standard practices. Proactively escalates unresolved or critical issues with a thorou