# OpenAI APIs - Embedding

SGLang provides OpenAI-compatible APIs to enable a smooth transition from OpenAI services to self-hosted local models.
A complete reference for the API is available in the [OpenAI API Reference](https://platform.openai.com/docs/guides/embeddings).

This tutorial covers the embedding APIs for embedding models, such as  
- [intfloat/e5-mistral-7b-instruct](https://huggingface.co/intfloat/e5-mistral-7b-instruct)  
- [Alibaba-NLP/gte-Qwen2-7B-instruct](https://huggingface.co/Alibaba-NLP/gte-Qwen2-7B-instruct)  


## Launch A Server

Launch the server in your terminal and wait for it to initialize. Remember to add `--is-embedding` to the command.

In [1]:
from sglang.utils import (
    execute_shell_command,
    wait_for_server,
    terminate_process,
    print_highlight,
)

embedding_process = execute_shell_command(
    """
python -m sglang.launch_server --model-path Alibaba-NLP/gte-Qwen2-7B-instruct \
    --port 30000 --host 0.0.0.0 --is-embedding
"""
)

wait_for_server("http://localhost:30000")

[2025-01-03 02:31:02] server_args=ServerArgs(model_path='Alibaba-NLP/gte-Qwen2-7B-instruct', tokenizer_path='Alibaba-NLP/gte-Qwen2-7B-instruct', tokenizer_mode='auto', load_format='auto', trust_remote_code=False, dtype='auto', kv_cache_dtype='auto', quantization=None, context_length=None, device='cuda', served_model_name='Alibaba-NLP/gte-Qwen2-7B-instruct', chat_template=None, is_embedding=True, revision=None, skip_tokenizer_init=False, return_token_ids=False, host='0.0.0.0', port=30000, mem_fraction_static=0.88, max_running_requests=None, max_total_tokens=None, chunked_prefill_size=8192, max_prefill_tokens=16384, schedule_policy='lpm', schedule_conservativeness=1.0, cpu_offload_gb=0, prefill_only_one_req=False, tp_size=1, stream_interval=1, random_seed=606729663, constrained_json_whitespace_pattern=None, watchdog_timeout=300, download_dir=None, base_gpu_id=0, log_level='info', log_level_http=None, log_requests=False, show_time_cost=False, enable_metrics=False, decode_log_interval=40, 

[2025-01-03 02:31:07] Downcasting torch.float32 to torch.float16.


[2025-01-03 02:31:15 TP0] Downcasting torch.float32 to torch.float16.


[2025-01-03 02:31:15 TP0] Overlap scheduler is disabled for embedding models.
[2025-01-03 02:31:15 TP0] Downcasting torch.float32 to torch.float16.
[2025-01-03 02:31:15 TP0] Init torch distributed begin.


[2025-01-03 02:31:16 TP0] Load weight begin. avail mem=78.81 GB
[2025-01-03 02:31:16 TP0] Ignore import error when loading sglang.srt.models.grok. unsupported operand type(s) for |: 'type' and 'NoneType'


[2025-01-03 02:31:16 TP0] Using model weights format ['*.safetensors']


Loading safetensors checkpoint shards:   0% Completed | 0/7 [00:00<?, ?it/s]


Loading safetensors checkpoint shards:  14% Completed | 1/7 [00:00<00:04,  1.31it/s]


Loading safetensors checkpoint shards:  29% Completed | 2/7 [00:01<00:05,  1.04s/it]


Loading safetensors checkpoint shards:  43% Completed | 3/7 [00:03<00:05,  1.36s/it]


Loading safetensors checkpoint shards:  57% Completed | 4/7 [00:05<00:04,  1.55s/it]


Loading safetensors checkpoint shards:  71% Completed | 5/7 [00:07<00:03,  1.65s/it]


Loading safetensors checkpoint shards:  86% Completed | 6/7 [00:09<00:01,  1.68s/it]


Loading safetensors checkpoint shards: 100% Completed | 7/7 [00:11<00:00,  1.75s/it]
Loading safetensors checkpoint shards: 100% Completed | 7/7 [00:11<00:00,  1.58s/it]

[2025-01-03 02:31:28 TP0] Load weight end. type=Qwen2ForCausalLM, dtype=torch.float16, avail mem=64.40 GB
[2025-01-03 02:31:28 TP0] Memory pool end. avail mem=7.42 GB


[2025-01-03 02:31:28 TP0] max_total_num_tokens=1028801, max_prefill_tokens=16384, max_running_requests=4019, context_len=131072
[2025-01-03 02:31:28] INFO:     Started server process [4034424]
[2025-01-03 02:31:28] INFO:     Waiting for application startup.
[2025-01-03 02:31:28] INFO:     Application startup complete.
[2025-01-03 02:31:28] INFO:     Uvicorn running on http://0.0.0.0:30000 (Press CTRL+C to quit)


[2025-01-03 02:31:29] INFO:     127.0.0.1:43060 - "GET /v1/models HTTP/1.1" 200 OK


[2025-01-03 02:31:29] INFO:     127.0.0.1:43076 - "GET /get_model_info HTTP/1.1" 200 OK
[2025-01-03 02:31:29 TP0] Prefill batch. #new-seq: 1, #new-token: 6, #cached-token: 0, cache hit rate: 0.00%, token usage: 0.00, #running-req: 0, #queue-req: 0


[2025-01-03 02:31:30] INFO:     127.0.0.1:43086 - "POST /encode HTTP/1.1" 200 OK
[2025-01-03 02:31:30] The server is fired up and ready to roll!


## Using cURL

In [2]:
import subprocess, json

text = "Once upon a time"

curl_text = f"""curl -s http://localhost:30000/v1/embeddings \
  -d '{{"model": "Alibaba-NLP/gte-Qwen2-7B-instruct", "input": "{text}"}}'"""

text_embedding = json.loads(subprocess.check_output(curl_text, shell=True))["data"][0][
    "embedding"
]

print_highlight(f"Text embedding (first 10): {text_embedding[:10]}")

[2025-01-03 02:31:34 TP0] Prefill batch. #new-seq: 1, #new-token: 4, #cached-token: 0, cache hit rate: 0.00%, token usage: 0.00, #running-req: 0, #queue-req: 0
[2025-01-03 02:31:34] INFO:     127.0.0.1:43088 - "POST /v1/embeddings HTTP/1.1" 200 OK


## Using Python Requests

In [3]:
import requests

text = "Once upon a time"

response = requests.post(
    "http://localhost:30000/v1/embeddings",
    json={"model": "Alibaba-NLP/gte-Qwen2-7B-instruct", "input": text},
)

text_embedding = response.json()["data"][0]["embedding"]

print_highlight(f"Text embedding (first 10): {text_embedding[:10]}")

[2025-01-03 02:31:34 TP0] Prefill batch. #new-seq: 1, #new-token: 1, #cached-token: 3, cache hit rate: 21.43%, token usage: 0.00, #running-req: 0, #queue-req: 0
[2025-01-03 02:31:34] INFO:     127.0.0.1:43098 - "POST /v1/embeddings HTTP/1.1" 200 OK


## Using OpenAI Python Client

In [4]:
import openai

client = openai.Client(base_url="http://127.0.0.1:30000/v1", api_key="None")

# Text embedding example
response = client.embeddings.create(
    model="Alibaba-NLP/gte-Qwen2-7B-instruct",
    input=text,
)

embedding = response.data[0].embedding[:10]
print_highlight(f"Text embedding (first 10): {embedding}")

[2025-01-03 02:31:34 TP0] Prefill batch. #new-seq: 1, #new-token: 1, #cached-token: 3, cache hit rate: 33.33%, token usage: 0.00, #running-req: 0, #queue-req: 0
[2025-01-03 02:31:34] INFO:     127.0.0.1:43102 - "POST /v1/embeddings HTTP/1.1" 200 OK


## Using Input IDs

SGLang also supports `input_ids` as input to get the embedding.

In [5]:
import json
import os
from transformers import AutoTokenizer

os.environ["TOKENIZERS_PARALLELISM"] = "false"

tokenizer = AutoTokenizer.from_pretrained("Alibaba-NLP/gte-Qwen2-7B-instruct")
input_ids = tokenizer.encode(text)

curl_ids = f"""curl -s http://localhost:30000/v1/embeddings \
  -d '{{"model": "Alibaba-NLP/gte-Qwen2-7B-instruct", "input": {json.dumps(input_ids)}}}'"""

input_ids_embedding = json.loads(subprocess.check_output(curl_ids, shell=True))["data"][
    0
]["embedding"]

print_highlight(f"Input IDs embedding (first 10): {input_ids_embedding[:10]}")

[2025-01-03 02:31:41 TP0] Prefill batch. #new-seq: 1, #new-token: 1, #cached-token: 3, cache hit rate: 40.91%, token usage: 0.00, #running-req: 0, #queue-req: 0
[2025-01-03 02:31:41] INFO:     127.0.0.1:35590 - "POST /v1/embeddings HTTP/1.1" 200 OK


In [6]:
terminate_process(embedding_process)