Founding Data Engineer

San Francisco, California

  Data Engineering

0

Permanent

Our client, a growing AI and Data organization, is hiring a Founding Data Engineer to join the team. The successful candidate will help transform enterprise data into actionable intelligence by building the backend systems, data pipelines, connector frameworks and graph-based knowledge models that power their Agentic AI applications.

Responsibilities

  • Build reliable, scalable data ingestion and transformation pipelines across structured, semi-structured and unstructured data sources.

  • Develop and maintain connector frameworks that integrate enterprise systems, including ERPs, PLMs, CRMs, legacy platforms, email, Excel, documents and web APIs.

  • Design and enhance the Data Fabric layer, including knowledge graphs enriched with ontologies, metadata and relationships.

  • Prepare, enrich, normalize and vectorize enterprise data for AI and LLM applications, including RAG, summarization and intelligent alerting.

  • Establish and maintain data contracts, access layers, lineage, governance and data quality processes.

  • Develop secure APIs that enable services, AI agents and users to access and query enriched semantic data.

  • Partner with ML and LLM teams to provide high-quality enterprise data for model training, tuning and AI applications.

  • Build robust systems that support schema evolution, observability, scalability and reliable production operations.

Skillset

  • Minimum of 5 years of experience designing, building and operating large-scale data infrastructure in production environments.

  • Strong hands-on experience with data ingestion frameworks such as Kafka, Airbyte, Meltano or Fivetran, alongside orchestration tools such as Airflow, Dagster or Prefect.

  • Proven experience working with structured, semi-structured, and unstructured data, including PDFs, Excel, emails, logs, CSVs and web APIs.

  • Experience with columnar data stores, object storage and modern lakehouse technologies such as Iceberg, Delta or Parquet.

  • Strong understanding of knowledge graphs and semantic modelling, with experience using platforms such as Neo4j, RDF, Gremlin or PuppyGraph.

  • Experience designing and developing GraphQL and REST APIs, as well as scalable and developer-friendly data access layers.

  • Strong knowledge of data governance principles, including RBAC, ABAC, data contracts, lineage and data quality.

  • Experience with vector databases and embedding pipelines, including technologies, such as Weaviate, Qdrant or Pinecone, is highly desirable.

  • Familiarity with enterprise connector ecosystems, ontology versioning, graph diffing and semantic schema alignment is an advantage.

  • Exposure to RAG pipelines, LLM fine-tuning, Linked Data, W3C standards or modern data fabric architectures is a plus.

Benefits

  • Salary: $150k – $200k.

  • Equity.