Together AI integration
Mask personal data before it reaches Together AI
- Every text field in the request body: user and system messages, tool and function-call arguments, and tool results
- Embeddings inputs, masked deterministically
- Streaming responses
- Base64, hex, and percent-encoded text inside the request is decoded and checked too
Together AI hosts open-weight models behind one API, so it is often where teams experiment, and experiments are where real customer data tends to slip in early.
Putting Maskflare in front of Together AI from the start means the same rules apply to experiments and production: sensitive values are masked before they leave, and embeddings stay consistent.
Set up Together AI with Maskflare
Requests go to /v1/together on your Maskflare gateway instead of the provider.
from openai import OpenAI
client = OpenAI(
base_url="https://YOUR_GATEWAY_HOST/v1/together",
api_key="mf_live_...", # a Maskflare key
)
reply = client.chat.completions.create(
model="meta-llama/Llama-3.3-70B-Instruct-Turbo",
messages=[{"role": "user", "content": "Summarise feedback from tom.berg@example.se"}],
)What gets masked
- Every text field in the request body: user and system messages, tool and function-call arguments, and tool results
- Embeddings inputs, masked deterministically
- Streaming responses
- Base64, hex, and percent-encoded text inside the request is decoded and checked too
Good to know
- A rule set to block stops the request before it reaches the provider and returns HTTP 403 with the matching rule keys, never the values.
- Store the provider key once in the Maskflare console, or keep sending it per request in pass-through mode.
- No code change option: the Maskflare forward proxy inspects traffic to api.together.xyz with the same rules.
Frequently asked questions
Does it work for every model on Together AI?
Yes. Masking applies to the request body regardless of model, for every path under /v1/together.
Can we use masked embeddings for retrieval?
Yes. Embeddings inputs are masked deterministically, so the same value always produces the same token and search over masked text stays consistent.
Can we start by only monitoring?
Yes. Set rules to alert: matches are counted in the traffic log while requests pass unchanged, so you can measure exposure before masking.
Other providers
See it on your own data
Book a 30-minute demo.
Bring your hardest prompts.
We'll show detection, masking, and restoration on your providers and data types,
and how it fits your stack.