
I will spend 60 minutes hands-on with you to get an OSS Large Language model from HuggingFace working in your GKE cluster.
There is no guarantee that we will be able to get it working within 60 minutes in your environment.
Example models that I've deployed succesfully on GKE using L4 GPUs: Falcon-7b or Falcon-40b.