OmDev Logo
GetYourJob
0
Publié aujourd'hui

INTERNSHIP - Data engineering & AI-ready software architecture

Entreprise
LesaffreRecrute en direct
Localisation
Marcq-En-Barœul, Hauts-de-France, France
Sur site
Type de contrat
Stage
Niveau
Junior

Description du poste

Company Description

About Lesaffre

Major global fermentation player for more than a century, Lesaffre, with more than 3 billion euros in revenue, has a presence on all continents, with 11,000 employees from over 90 nationalities. Using our expertise and diversity, we collaborate with clients, partners, and scientists to find increasingly relevant responses to the needs of nutrition, health, naturalness, and respect for our environment. Thus, every day, we explore and reveal the infinite potential of microorganisms.

Feeding 9 billion people in 2050 in a healthy way while using the planet's resources in the most efficient manner is a major and unprecedented challenge. We believe that fermentation is one of the most promising answers to this challenge.

 

Lesaffre - Working together to better nourish and protect the planet.

Environment

Lesaffre is a leader in bioengineering & bioprocesses and places RD&I at the heart of its success: more than 850 experts worldwide explore the potential of microorganisms and natural fermentation to serve human, plant, and animal needs. 

You will join this research community at the interface of two RD&I teams: NMH (Nutrition, Microbiota & Health), which runs experimental science (microbiological assays, imaging, microbiota work), and Biodata, focusing on data science and bioinformatics team that builds the computational tools used by our scientists. You will be co-mentored by an NMH Scientist and a BioData DevOps engineer.

Job Description

The scientific problem

Every week, NMH experiments produce large volumes of heterogeneous measurements: plate-reader kinetics, microbiological assay readouts, imaging data, each in the native format of the instrument that produced it. Today, turning that raw output into an analysable result is largely manual:

  • Experimental time is lost to data wrangling rather than to designing and running the next experiment.
  • Results are hard to compare across campaigns, because the same biological variable is recorded differently depending on the instrument, the operator, and the date.
  • Reprocessing a past experiment is difficult, which limits reproducibility and prevents the accumulated data from being reused to generate new hypotheses.

Internship goal: build a web interface in streamlit, python layer that takes raw instrument output to a clean, described, queryable experimental result and make that result directly consumable by modern AI architectures (LLMs, RAG pipelines, autonomous agents), so that a scientist can interrogate their own data in natural language.

Key responsibilities

  • Instrument data acquisition & parsing: Design automated extraction modules capable of ingesting raw experimental data streams from various laboratory instruments and heterogeneous file formats, ensuring cleaning, typing, and quality control.
  • Data standardization & schema modeling: Define and implement strict, unified data schemas (metadata, type validation, JSON serialization formats) to ensure data interoperability and reproducibility.
  • Architecture: Structure datasets and develop tools/function calling endpoints, enabling AI agents or LLM/RAG pipelines to autonomously query, analyze, and manipulate standardized datasets.
  • API & internal tooling: Expose processing services via REST APIs and build lightweight interfaces (e.g., Streamlit) enabling scientific teams to visualize, validate, and query data independently.
  • Apply Biodata’s software engineering best practices: containerization (Docker), version control (Git), unit/integration testing, and technical documentation.

Expected deliverables

By the end of the internship you will deliver a deployed, operational standardization pipeline for at least one NMH data family; a documented experimental data schema; an LLM‑queryable access layer; and a lightweight interface used by the scientific team. Success will be demonstrated via a live demo, automated tests, deployment instructions, a short technical report, and readiness for inclusion in the supported scientific campaign.

Exigences du poste

Compétences requises

  • Git
  • JSON
  • Streamlit
  • Docker
  • Python

Un plus

  • DevOps
  • Next.js
  • Makefile
  • LLM
  • RAG
  • API Development
  • REST API
  • QA Testing
  • Agentic AI / Agents IA
  • Data Engineering
  • Data Science
  • Software Architecture

Toutes les offres Python à Lille

Plan d'action

Un plan personnalisé pour postuler intelligemment à cette offre.

À propos de l'entreprise

LesaffreRecrute en direct
Voir toutes les offres de Lesaffre

Publié par

Recruteur
Recruteur

Intéressé par cette offre ?

Cliquez sur "Postuler" pour accéder à l'offre.