Founding Data Engineer
Our client, a growing AI and Data organization, is hiring a Founding Data Engineer to join the team. The successful candidate will help transform enterprise data into actionable intelligence by building the backend systems, data pipelines, connector frameworks and graph-based knowledge models that power their Agentic AI applications.
Responsibilities
-
Build reliable, scalable data ingestion and transformation pipelines across structured, semi-structured and unstructured data sources.
-
Develop and maintain connector frameworks that integrate enterprise systems, including ERPs, PLMs, CRMs, legacy platforms, email, Excel, documents and web APIs.
-
Design and enhance the Data Fabric layer, including knowledge graphs enriched with ontologies, metadata and relationships.
-
Prepare, enrich, normalize and vectorize enterprise data for AI and LLM applications, including RAG, summarization and intelligent alerting.
-
Establish and maintain data contracts, access layers, lineage, governance and data quality processes.
-
Develop secure APIs that enable services, AI agents and users to access and query enriched semantic data.
-
Partner with ML and LLM teams to provide high-quality enterprise data for model training, tuning and AI applications.
-
Build robust systems that support schema evolution, observability, scalability and reliable production operations.
Skillset
-
Minimum of 5 years of experience designing, building and operating large-scale data infrastructure in production environments.
-
Strong hands-on experience with data ingestion frameworks such as Kafka, Airbyte, Meltano or Fivetran, alongside orchestration tools such as Airflow, Dagster or Prefect.
-
Proven experience working with structured, semi-structured, and unstructured data, including PDFs, Excel, emails, logs, CSVs and web APIs.
-
Experience with columnar data stores, object storage and modern lakehouse technologies such as Iceberg, Delta or Parquet.
-
Strong understanding of knowledge graphs and semantic modelling, with experience using platforms such as Neo4j, RDF, Gremlin or PuppyGraph.
-
Experience designing and developing GraphQL and REST APIs, as well as scalable and developer-friendly data access layers.
-
Strong knowledge of data governance principles, including RBAC, ABAC, data contracts, lineage and data quality.
-
Experience with vector databases and embedding pipelines, including technologies, such as Weaviate, Qdrant or Pinecone, is highly desirable.
-
Familiarity with enterprise connector ecosystems, ontology versioning, graph diffing and semantic schema alignment is an advantage.
-
Exposure to RAG pipelines, LLM fine-tuning, Linked Data, W3C standards or modern data fabric architectures is a plus.
Benefits
-
Salary: $150k – $200k.
-
Equity.
SHARE JOB
related jobs