Basé sur 170 offres pour ce poste (tous niveaux, Île-de-France, 3 dernières semaines). Fourchette habituelle 450€/j–595€/j, médiane 530€/j. Cette offre (525€/j) est dans la fourchette.
Our client is seeking DevOps and Site Reliability Engineering expertise to help build and operate a secure, regulated, and highly available platform for tokenised financial instruments.
They are developing a new Digital Assets capability to support tokenisation, distributed ledger technology, and the next generation of financial market infrastructure.
They need expertise in infrastructure-as-code, Kubernetes, container platforms, and CI/CD pipelines to establish scalable foundations and enable controlled, zero-downtime deployments.
They require comprehensive observability across metrics, logs, traces, alerting, service level objectives, and error budgets to support a live 24/7 service. We are also looking for operational guidance covering on-call management, incident response, actionable runbooks, postmortems, and platform stability during rapid product and technology changes.
The engagement should incorporate secure and auditable infrastructure practices suitable for regulated workloads, including collaboration with security stakeholders and disciplined controls around AI-assisted engineering. We additionally seek expertise in applying AIOps for anomaly detection, alert triage, incident summarisation, and the monitoring, operation, and rollback of in-product AI/ML services.
Establish a resilient operational foundation for a regulated digital asset platform and support reliable service delivery at scale.
Design and implement maintainable infrastructure-as-code, including Terraform-based provisioning and version-controlled configuration.
Build and optimise secure CI/CD pipelines supporting continuous integration, continuous deployment, automated validation, approvals, auditability, and zero-downtime releases.
Architect and operate containerised workloads using Kubernetes and related orchestration technologies.
Implement an end-to-end observability solution covering metrics, centralised logging, distributed tracing, dashboards, alerting, and operational health indicators.
Define and operationalise service level objectives, service level indicators, and error-budget practices aligned with platform reliability requirements.
Establish an effective 24/7 on-call operating model, including escalation procedures, incident workflows, ownership standards, and service continuity practices.
Produce and maintain clear, actionable runbooks for recurring operational tasks, failure scenarios, deployments, recoveries, and incident response.
Strengthen infrastructure security through hardening, least-privilege access, secrets management, vulnerability remediation, compliance controls, and auditable change management.
Apply AI-assisted engineering tools to infrastructure-as-code and pipeline authoring with documented human review and deployment approval controls.
Implement AIOps capabilities for anomaly detection, alert triage, incident summarisation, and postmortem drafting.
Deploy, monitor, operate, and safely roll back in-product AI/ML services, including performance, availability, and model operational monitoring.
Deliver technical documentation, architecture decisions, operational procedures, reliability metrics, and recommendations for continuous improvement.
Support platform stability and availability through periods of rapid product evolution while maintaining regulated-workload standards and zero-downtime objectives.
Un plan personnalisé pour postuler intelligemment à cette offre.
Cliquez sur "Postuler" pour accéder à l'offre.