Job Description
Infrastructure Solution Architect (DR for data centers. Oracle Platform, Automating Failover, networking and security)
- Rate TBD
- Location Canada
- Type of project
- Duration
- Education required N/A
- Years of experience N/A
- Type of employment N/A
- Area of Specialization N/A
- Languages required N/A Workhoppers Home
- Description
Infrastructure Solution Architect (DR for data centers. Oracle
Platform, Automating Failover, networking and security, storage )
Work Location: Remote Work in Canada (flexibility to work EST)
Contract Duration: 9 months, renewable
We are looking for a systems thinker who treats infrastructure as a
single, resilient system — from Oracle and data centers to networks,
load balancing and security - and who can automate, validate and
continuously improve disaster recovery (DR) and failover readiness.
This role partners across engineering, operations, security and
product teams to design resilient architectures, implement automated
failover testing, and ensure predictable recovery outcomes for
mission-critical services.
Key responsibilities
Own DR and resilience strategy for multi-tier applications and Oracle
estates: define RTO/RPO targets, architecture patterns (active/active,
active/passive, warm/cold), replication approaches and acceptance
criteria.
Design and implement automated failover orchestration and validation
pipelines using IaC, orchestration and runbook automation tools (e.g.,
Terraform, Ansible, Jenkins/Rundeck, custom scripts).
Create repeatable, scheduled failover test workflows that exercise
full-stack recovery (compute, storage, networking, DNS/BGP, load
balancers, firewall policies, and application state).
Build synthetic transaction suites and observability checks to verify
service correctness during and after failover; automate rollup of
pass/fail metrics and evidence for audits.
Lead live and simulated DR exercises (tabletop and full cutover),
coordinate cross-functional playbooks, and manage communications and
escalation paths during tests and incidents.
Architect resilient network and routing designs to support automated
cutover and traffic steering (circuit failover, BGP announcements, DNS
TTL strategies, load balancer health probing).
Integrate cyber controls into DR plans: ensure segmented recovery
zones, secure key/material handling, forensic logging, and validated
access controls during failover.
Define and maintain DR runbooks, recovery automation code, checklists,
and post-test action items; ensure runbooks are versioned and
testable.
Measure, report, and improve DR maturity: track test frequency,
success rates, RTO/RPO attainment, MTTD/MTTR improvements and risk
reduction.
Mentor teams on resilient application design (stateless vs. stateful
components, session management, data replication consistency) and
incorporate resilience patterns into CI/CD pipelines.
Participate in incident postmortems and convert findings into
automated safeguards and improved test coverage.
Experience & technical skills
Hands-on experience designing and operating DR for enterprise-scale
data centers and hybrid environments with measurable RTO/RPO results.
Deep Oracle platform knowledge (replication, Data Guard, RMAN,
backup/restore strategies) and experience validating database recovery
in automated failovers.
Proven skill automating failover workflows and tests using
orchestration/IaC tools (Terraform, Ansible, Rundeck, Jenkins, or
comparable tooling) and scripting (Python, Bash, PowerShell).
Strong networking and routing expertise: BGP, DNS failover strategies,
load balancer (F5/HAProxy/cloud LB) configurations and circuit
redundancy planning.
Experience with storage replication (SAN/NAS/async-sync), consistency
models and impact on application recovery.
Observability and verification of tooling: metrics, logs, tracing,
synthetic transactions; ability to automate health checks and roll-up
reporting during tests.
Security and compliance awareness as applied to DR: key management,
access controls, hardened recovery environments and audit evidence
generation.
Practical experience running tabletop and live DR exercises,
documenting outcomes, and driving remediation through automation.
Excellent communication skills for cross-team coordination during
planned tests and unplanned incidents.
Preferred / nice-to-have
Certifications: Oracle Certified Professional, CCNP, CISSP, SRE/DevOps
certifications, cloud provider DR certs.
Experience with chaos engineering practices and tools (Chaos Monkey,
Gremlin) applied to infrastructure-level resilience testing.
Familiarity with carrier management, circuit provisioning, and SLA
negotiation for DR scenarios.
Prior experience in regulated industries where DR testing evidence is
required for audits.
Candidate profile / behavioral traits
Systems-level thinker who balances architectural rigor with
operational pragmatism.
Proactive collaborator and facilitator — we rely on this role to
lead cross-functional DR exercises and embed automated safeguards.
Comfortable for both hands-on (writing automation, running tests) and
at the leadership level (designing strategy, reporting metrics).
Data-driven, with an emphasis on measurable improvements and
continuous validation.
Calm under pressure, decisive during incident response and committed
to blameless postmortems and remediation.
-May 14, 2026