AI architecture
We build RAG pipelines with technologies such as pgvector and Pinecone, and design agent workflows with LangGraph. We make sources and retrieval traceable and use guardrails to reduce unsupported responses.
We build RAG pipelines with technologies such as pgvector and Pinecone, and design agent workflows with LangGraph. We make sources and retrieval traceable and use guardrails to reduce unsupported responses.
Every trained model can be traced back to the data, configuration and experiments behind it. We build training and serving pipelines with Kubeflow, define infrastructure with Terraform and track experiments in MLflow. Environments are reproducible from code.
We run Kubernetes workloads on EKS and GKE and give teams visibility into their infrastructure and service costs, with cost allocation where the architecture supports it.
We design systems to tolerate failures within defined failure domains. Depending on the workload, this includes multi-AZ deployments, health checks, automated failover and self-healing Kubernetes workloads.
Infrastructure scales with actual demand instead of fixed capacity. We use HPA for pod scaling and Karpenter for node provisioning where Kubernetes workloads benefit from dynamic capacity.
We design AWS environments for workloads with demanding security, audit and operational requirements. Depending on the workload, this can include SageMaker for ML workloads, Lambda for event-driven processing, S3 for data storage and CloudWatch for monitoring and alerting.
Your team can see what is running, what changed and where problems are occurring. We build CI/CD pipelines with GitHub Actions and Argo CD, monitor systems with Prometheus and Grafana, and integrate alerting with tools such as PagerDuty.
ML models are part of an application, not an isolated component. We expose models through clearly defined APIs and design for failures with structured logging, timeouts, health checks and controlled fallback paths.
Security controls are built into the delivery process from the beginning. We work with IAM policies, network isolation, KMS encryption and audit logging, with controls designed to support requirements such as GDPR and SOC 2.
We build computer vision systems for real-world production environments, from dataset creation through deployment and monitoring. Each system is designed around its operating environment: lighting, camera position, available compute and network constraints.
We build inference pipelines with ONNX Runtime and TensorRT and serve them through production APIs with health checks, structured logging and monitoring. Where required, prediction inputs and outputs can be made traceable for debugging and audit.
We look at the data before changing the model. We inspect and label datasets with tools such as Label Studio and focus on the cases that actually cause errors. In many production systems, improving the data is more valuable than moving to a larger model.
We train detection and other vision models through reproducible experiments and controlled deployment processes. Models can be released through staged or canary rollouts, with production feedback used to guide subsequent training.
When latency, connectivity or data locality matters, we deploy vision models to edge hardware such as NVIDIA Jetson. Heavier workloads can run on GPU infrastructure in the cloud. We measure latency, throughput and cost before deployment, not after.
When model performance changes, you need to know what changed with it. We version datasets and training configurations with tools such as DVC and Git LFS, so training runs can be compared and changes traced over time.
Production data changes. We monitor prediction confidence and operational signals, use tools such as Evidently to identify potential data drift, and route low-confidence or uncertain cases to a person for review when appropriate.
Kaya is Glotech's own product. It reads the tower light on each machine through the cameras your factory already has, and shows how many hours each machine actually ran.
Kaya reads each machine's tower light as green, amber, red or off. It turns that into hours per machine, per day and per shift. When the light is amber and nobody is at the machine, that time is shown separately and not counted.
Kaya uses the cameras already on the factory floor. You don't add sensors, wiring or a PLC connection. If your camera recorder (NVR/DVR) can only be reached from inside the factory, one small computer sits on your network. It touches no machine. Make and age don't matter. Kaya works on any machine with a tower light: CNC machines, presses, injection moulding machines, packaging lines.
Open any hour in a report to see the camera frames it was measured from, the light colour Kaya read and how sure it was. If the camera could not see the light, that time is marked as no data, never as zero.
Glotech is an AI infrastructure and MLOps consultancy. We take AI systems from development into production.
No. AWS is a large part of our work, but we also work on Google Cloud. Most of what we build runs on open-source tools that work on either cloud.
No. Kaya is a product Glotech built to measure how long machines run. It is sold separately from our consultancy work.