Fabrion Jobs

Data Engineer (Founding Team)

Fabrion

Data Engineer (Founding Team)

Reposted 18 Days Ago

In-Office or Remote

Hiring Remotely in CA

Senior level

In-Office or Remote

Hiring Remotely in CA

Senior level

Build and operate scalable data ingestion, transformation, and connector frameworks; design and maintain a knowledge-graph-based data fabric; normalize and vectorize enterprise data for LLM/AI workflows; implement governance, lineage, access controls, and secure APIs to serve ML/agent pipelines.

The summary above was generated by AI

Data/ETL Engineer (Founding Team)

Location: San Francisco Bay Area

Type: Full-Time

Compensation: Competitive salary + early-stage equity

Backed by 8VC, we're building a world-class team to tackle one of the industry’s most critical infrastructure problems.

About the Role

We’re building a multi-tenant, AI-native platform where enterprise data becomes actionable through semantic enrichment, intelligent agents, and governed interoperability. At the heart of this architecture lies our Data Fabric — an intelligent, governed layer that turns fragmented and siloed data into a connected ontology ready for model training, vector search, and insight-to-action workflows.

We're looking for engineers who enjoy hard data problems at scale: messy unstructured data, schema drift, multi-source joins, security models, and AI-ready semantic enrichment. You’ll build the backend systems, data pipelines, connector frameworks, and graph-based knowledge models that fuel agentic applications.

If you've worked on streaming unstructured pipelines, built connectors into ugly legacy systems, or mapped knowledge graphs that scale — this role will feel like home.

Responsibilities

Build highly reliable, scalable data ingestion and transformation pipelines across structured, semi-structured, and unstructured data sources
Develop and maintain a connector framework for ingesting from enterprise systems (ERPs, PLMs, CRMs, legacy data stores, email, Excel, docs, etc.)
Design and maintain the data fabric layer — including a knowledge graph (Neo4j or Puppygraph) enriched with ontologies, metadata, and relationships
Normalize and vectorize data for downstream AI/LLM workflows — enabling retrieval-augmented generation (RAG), summarization, and alerting
Create and manage data contracts, access layers, lineage, and governance mechanisms
Build and expose secure APIs for downstream services, agents, and users to query enriched semantic data
Collaborate with ML/LLM teams to feed high-quality enterprise data into model training and tuning pipelines

What We’re Looking For

Core Experience:

5+ years building large-scale data infrastructure in production environments
Deep experience with ingestion frameworks (Kafka, Airbyte, Meltano, Fivetran) and data pipeline orchestration (Airflow, Dagster, Prefect)
Comfortable processing unstructured data formats: PDFs, Excel, emails, logs, CSVs, web APIs
Experience working with columnar stores, object storage, and lakehouse formats (Iceberg, Delta, Parquet)
Strong background in knowledge graphs or semantic modeling (e.g. Neo4j, RDF, Gremlin, Puppygraph)
Familiarity with GraphQL, RESTful APIs, and designing developer-friendly data access layers
Experience implementing data governance: RBAC, ABAC, data contracts, lineage, data quality checks

Mindset & Culture Fit:

You’re a system thinker: you want to model the real world, not just process it
Comfortable navigating ambiguous data models and building from scratch
Passionate about enabling AI systems with real-world, messy enterprise data
Pragmatic about scalability, observability, and schema evolution
Value autonomy, high trust, and meaningful ownership over infrastructure

Bonus Skills

Prior work with vector DBs (e.g. Weaviate, Qdrant, Pinecone) and embedding pipelines
Experience building or contributing to enterprise connector ecosystems
Knowledge of ontology versioning, graph diffing, or semantic schema alignment
Familiarity with data fabric patterns (e.g. Palantir Ontology, Linked Data, W3C standards)
Familiar with fine-tuning LLMs or enabling RAG pipelines using enterprise knowledge
Experience enforcing data access policy with tools like OPA, Keycloak, Snowflake row-level security

Why This Role Matters

Agents are only as smart as the data they operate on. This role builds the foundation — the semantic, governed, connected substrate — that makes autonomous decision-making and agent action possible. From factory ERP records to geopolitical news alerts, the data fabric unifies it all.

If you're excited to tame complexity, unify chaos, and power intelligent systems with trusted data — we’d love to hear from you.

Similar Jobs

Prolific

Human Data Quality Engineer (Founding Team)

3 Days Ago

Remote

Canada

Senior level

Artificial Intelligence

Design and own end-to-end data quality systems for managed AI data studies: define rubrics, sampling plans, automated checks, launch gates, drift detection, dashboards, and calibration. Investigate integrity issues, run root-cause analysis, build automation in Python/SQL, and partner with Product, Engineering, and Operations to embed quality controls, train reviewers, and scale repeatable quality playbooks across programs.

Top Skills: PythonSQL

Dynatrace

Account Executive

6 Hours Ago

Remote or Hybrid

Calgary, AB, CAN

Senior level

Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Big Data Analytics • Automation

Hunt and close net-new enterprise accounts across Alberta, building executive relationships and value-based business cases. Drive pipeline via prospecting, partners, and events; lead complex sales cycles with Sales Engineers and Customer Success to deliver observability, security, cloud modernization, and AI-driven outcomes.

Top Skills: AiopsAWSAzureCloud MigrationCybersecurityDevOpsDynatraceGCPObservabilityPlatform Engineering

Applied Systems

Senior Network Engineer

9 Hours Ago

Remote or Hybrid

Canada

Senior level

Cloud • Insurance • Payments • Software • Business Intelligence • App development • Big Data Analytics

Design, deploy, and operate enterprise network infrastructure across data centers and cloud (GCP, AWS, Azure). Manage VPCs, VPNs, BGP, F5 BIG-IP, hybrid connectivity, and observability. Build IaC and automation with Terraform, GitLab CI/CD, and Ansible. Maintain monitoring, documentation, and participate in on-call support for critical incidents.

Top Skills: AnsibleAWSAzureBgpCloud RouterCloudflareDnsF5 Big-IpFirewallsGCPGitlab Ci/CdHa VpnIp Address ManagementIpsec VpnLoad BalancingTerraformVpc

What you need to know about the Calgary Tech Scene

Employees can spend up to one-third of their life at work, so choosing the right company is crucial, not just for the job itself but for the company culture as well. While startups often offer dynamic culture and growth opportunities, large corporations provide benefits like career development and networking, especially appealing to recent graduates. Fortunately, Calgary stands out as a hub for both, recognized as one of Startup Genome's Top 100 Emerging Ecosystems, while also playing host to a number of multinational enterprises. In Calgary, job seekers can find a wide range of opportunities.