janitorbeta
newshelp
LoginRegister

Proxy Guide (Free)

Proxy Guide (Free)

4k

11k

by:@nalabyst

(Last Updated: August 5th 2026)

Use proxy. JLLM is not that good, and it can't handle many tokens. (I swear, it's addicting).

I tested a few proxies for my bots (all with free plans), and here is how you can use each one of them:



The way proxy works has made me rework this guide entirely many times up to now, so I decided to change my approach slightly. This guide will now be divided into a few segments, being a "General Guide" at the top — for an explanation of proxy implementation through any routers — and specific guides (such as "Openrouter Guide") below, with more summed up instructions on how to set up your proxy with recommended routers, and recommended models.

Because of that, since information is now a bit spread out, I will leave what I believe to be the best service and AI model here at the top.

Best (Free) Models:

TokenReply: minimaxai/minimax-m3

Nvidia Nim: z-ai/glm-5.2

Gemini: gemini-3.5-flash-lite

Openrouter: google/gemma-4-31b-it:free


I have also seen some people having trouble setting it up from this guide. Since the comments have a character limit, if you need any help feel free to send me a DM in Discord: nalabyst



General Guide (Janitor)

For any configuration you want to set up, you will always have to follow the same steps with slight differences. While in Janitor:

• Start a chat with any character

• Click on the purple button above saying "using janitor" or "using [something]" and a new tab will pop up

• Select "Proxy"

• Scroll down to Proxy Configurations and click in "+ New", and it will ask for a few settings:

° Name: you can put whatever you want in here

° Model: model ID (for example: "deepseek-free")

° Proxy URL: endpoint for the router you are using (for example: https://openrouter.ai/api/v1/chat/completions)

° API Key: the API Key that you will create in a specific router

(for all of these I am going to present more information in the specific guides below)

• There are also two sections for prompts, that are important to lead the LLM to behave as you wish:

° Global Prompt: I mostly use this for setting up the prompt, and I like the prompt from this guide

° Custom Prompt (for each model): I only use this when I notice that a certain model needs some fine tuning in the instructions, or for models like Gemini that can flag jailbreaks or NSFW content

• At last, click on "Save", refresh the page, and it's done!

Generation Settings:

The range in brackets represents the usual range for that setting across most models.

• Temperature [0.6 - 1.2]: how "creative" your model will be — higher values means it's more creative. Turning this up too high will make the AI say incoherent stuff. This varies a lot depending on the model you are using.

• Max Tokens [0 - 2k]: how many tokens the AI can generate as output (response). You can set it to 0 and it will be unlimited, and this is not that important, I leave mine at 1.5k just to avoid some glitches.
Note: If you are using "thinking" or "reasoning" models, set this to 0, or it will limit the amount of tokens it can use to generate it's thought process.

• Context Size [32k / 64k / 128k]: how many tokens the AI can take as input (context). For this one you will usually want to set it to the max that the model you are using can take. Most good models are in the range of 32k or 64k contexts, some even 128k and 256k. Always check the limit of the model you are using, and set the context size to be inside that limit.

Advanced:

If you don't want to mess too much with this, you can just set everything to 0 and Janitor will use the model's default. But here's a summary of what each setting is for:

• Top K [30 - 50]: the amount of words the AI has to pick from — the higher, the more words it will gain access to use.

• Top P [0.9 - 0.99]: the probability of the AI picking a more unique word, basically how much of the vocabulary from Top K it can really access.

• Repetition Penalty [0.8 - 1.3]: higher = makes the AI repeats less words in a message.
Note: If you are using "thinking" or "reasoning" models, set this to 0, or it will limit the repetition of words in the model's reasoning.

• Frequency Penalty [0.1 - 0.3]: higher = makes the AI repeats less words in the total chat.
Note: If you are using "thinking" or "reasoning" models, I also recommend setting this to 0. This is not as bad as the other two configurations with this note, but it still affects the model's reasoning over long chats.

Prefill:

You can enable it if you are getting errors, but leave it disabled otherwise.



TokenReply Guide

TokenReply is still my favorite model provider, with decent models (including unlimited ones), as well as availability for each model. Though, some free models seem to get really unstable with high usage.

TokenReply has a general rate limit of 40 requests per day for limited models, and a 3 requests per minute limit for unlimited ones.

Setup:

• Enter tokenreply.com

• Create an account and log in

Proxy URL:

Paste it in the "Proxy URL" field on Janitor (mentioned in the General Guide):

• Default: https://api.tokenreply.com/v1

• Modified (Sophia's Lorebay): (custom — more information in the "Sohpia's Lorebay Guide" below)

API Key:

Paste it in the "API Key" field on Janitor (mentioned in the General Guide):

• Go to "Console" at the top

• Copy your API Key and paste it on Janitor

Note: After you create your account, you can use any model labeled as "Free" in the model list without any hard limits. However, sometimes there are a $1 claim you can take, or you can use the daily check-ins to add $0.01 to your wallet — which allow for you to use the "Weekly Featured" models.

Models:

You can search for the available models on your own and test them, but here are my recommendations:

Free (Fully Unlimited):

• minimaxai/minimax-m3

• google/gemma-4-26b-a4b-it

• minimaxai/minimax-m2.7

• stepfun-ai/step-3.7-flash

• grok-4.3-high

Weekly Featured (Limited):

• gemini-2.5-flash

• deepseek-v4-pro

• deepseek-v4-flash-thinking

• kimi-k2.7

You can just paste the model IDs in the "Model" field on Janitor (mentioned in the General Guide).

Each one of the models you set up will have better settings that change drastically their performance (specially the Temperature). If you want you can search it up, or if you are too lazy just ask another AI, like Gemini or ChatGPT, what are the best parameters for each one of the models you find.

P.S.: If you are getting too many errors with a model using this router, check the model availability. If it's below 50%, you should probably switch to another model until it goes back up.



Nvidia Nim Guide

⚠️ Most models in Nvidia Nim seem to be giving errors at the moment ⚠️

I have said before I would only keep a max of 3 providers at the same time in this guide, but this one deserves it's place, and I don't want to remove Openrouter.

I recently found out about Nvidia Nim by somebody's comment on this post, and even though it has only a few models available, they are pretty good models with the only hard limit being 40 RPM (Requests Per Minute).

Setup:

• Enter https://developer.nvidia.com/nim

• Click on "Try APIs"

• Click on "Login" (even if you are signing up) and enter your email

Proxy URL:

For some reason, you cannot use their Proxy URL directly in Janitor or you will get a network error, so for this you will need to set it up with Sophia's Lorebay.

For that, check the "Sophia's Lorebay Guide" section at the bottom of this guide, and set it up with this URL: https://integrate.api.nvidia.com/v1/chat/completions

API Key:

I believe you will need to verify your phone number to get access to your API Key.

Paste it in the "API Key" field on Janitor (mentioned in the General Guide):

• Click on your avatar (usually top right)

• Click on "API Keys"

• Click on "Generate API Key"

• Give it a name of your choosing and select "Never Expire" for Expiration

• Copy the key and paste it on Janitor

Models:

You can search for the available models on your own and test them, but here are my recommendations:

• z-ai/glm-5.2

• deepseek-ai/deepseek-v4-pro

• nvidia/nemotron-3-ultra-550b-a55b

• thinkingmachines/inkling



Gemini Guide (SFW only)

I have always liked the Gemini models a lot when the quotas were more generous, but for a good time now it has been almost impossible to use, with hard limits and lots of errors.

Recently, I found a way to use Gemini with (practically) unlimited responses, and basically without errors. The only downside you need to be aware of is that Gemini is very restrictive on NSFW content, where even if you do jailbreak it, you might end up getting banned.

Also, a personal advice: do not create multiple accounts to bypass the quota limits. Besides it being against Google's TOS, I got banned for that before lol.

Setup:

• Enter aistudio.google.com

• Log in using your Google account

Proxy URL:

Paste it in the "Proxy URL" field on Janitor (mentioned in the General Guide):

• Enter GeminiForJanitor: gfjproxies.vu5eruz.workers.dev (this is not the Proxy URL, you need to open it)

• Copy the URL with the lowest bandwith (higher will give you more errors) and paste it on Janitor

You can also set up a custom URL for Sophia's Lorebay using the method in the "Sophia's Lorebay Guide" below

API Key:

Paste it in the "API Key" field on Janitor (mentioned in the General Guide):

• Click on the key symbol on the bottom left corner ("Get API key")

• Click on the button on the top right that says "Create API Key"

• Click on "Create key"

• Copy your API Key and paste it on Janitor

Models:

You can search for the available models on your own in here: ai.google.dev/gemini-api/docs/models and test them for yourself. However, as far as I know, you cannot access versions older than Gemini 3 through that URL.

Each model has a rate limit that you can check on "Rate Limit" in AI Studio. Generally, the limits are:

• 5 Requests per Day (RPD) for Pro models (e.g. gemini-3.1-pro)

• 20 Requests per Day (RPD) for Flash models (e.g. gemini-3.6-flash)

• 500 Requests per Day (RPD) for Flash Lite models (e.g. gemini-3.5-flash-lite)

Here are my recommendations:

• gemini-3.5-flash-lite

• gemini-3.1-flash-lite

• gemini-3.6-flash



Openrouter Guide

Openrouter is probably the most known model router available, being the most stable one. Recently, it has been in it's "dark ages", with not many good models.

Openrouter has a general limit rate of around 50 requests per day for free.

Setup:

• Enter openrouter.ai

• Create an account and log in

Proxy URL:

Paste it in the "Proxy URL" field on Janitor (mentioned in the General Guide):

• Default: https://openrouter.ai/api/v1/chat/completions

• Modified (Sophia's Lorebay): https://api.lorebary.com/openrouter

API Key:

Paste it in the "API Key" field on Janitor (mentioned in the General Guide):

• Go to you account settings (usually in the drop down menu by your icon)

• Go to "API Keys"

• Click on "Create"

• It will ask for a few settings, but they are not important — you can just set a name of your choosing and click "Create"

• Copy your key and paste it on Janitor

Models:

You can search for the available models on your own and test them, but here are my recommendations:

• google/gemma-4-31b-it:free

• nousresearch/hermes-3-llama-3.1-405b:free (I couldn't even test this because of errors 429, but I heard it's decent)

• nvidia/nemotron-3-ultra-550b-a55b:free

You can just paste the model IDs in the "Model" field on Janitor (mentioned in the General Guide).

Each one of the models you set up will have better settings that change drastically their performance (specially the Temperature). If you want you can search it up, or if you are too lazy just ask another AI, like Gemini or ChatGPT, what are the best parameters for each one of the models you find.



Other Routers

There are a few other router options that I tested and I think are decent. The process is the same:

• Create an account and log in

• Create and copy an API Key then paste it on Janitor

• Search the docs for the Proxy URL (and if it doesn't work try adding "/v1" or "/v1/chat/completions" at the end) and paste it on Janitor

• Search for models and paste the ID on Janitor


Other options I tried, if you want to check out for yourself, are:

• Electron Hub (model examples: Deepseek v4 Flash, Kimi K2.5 — it's actually very good too, but it's temporarily unavailable for free)

• NavyAI (model examples: Gemini v2.5 Pro — but has a very low token limitation per day)

• MeganovaAI (model examples: Sapphira-L3.3-70B-0.1 — decent but not as good as the others)



Sophia's Lorebay Guide

For any routers you can set them up to work with Sophia's Lorebay. This is a service made by the community that has many functions to improve your chats, the only downside I have heard is that it has some added restrictions for extreme NSFW content.

Setup:

• Go to lorebary.sophiamccarty.com

• Create an account and log in

• Hover over "Connection" and go to "Connect Proxy"

• Under the "Connect my current app" section, click on "Show me the setup"

• There are default URLs for popular routers (such as Openrouter) for you to pick from, but if you don't see the service you are using there, scroll to the side and click on "Other provider"

• Click on "Add Proxy" and set it up:

° Endpoint URL: paste the default Proxy URL for the router you are using

° Nickname: give it a name of your choosing

• Click on "Create & Test"

• Most of the time it will give an error, but just click on "Continue anyway". If you don't get an error, click on "Nice, continue"

• Copy the URL and click on "Got it"

• Paste it in the "Proxy URL" field on Janitor (mentioned in the General Guide)

Server Commands:

With Sophia's URL set up, you can use some "commands", which are just prompt injections that will give instructions to the AI to enhance specific behaviors. You can look at the list of full commands by going to "Proxy Portal" and then "Server Commands".

You can insert the commands in your custom prompt, in your chat memory, or the bot's definition. A few examples of commands I recommend:

<NOOMNISCIENCE>
Characters only know what they witnessed, were told, or logically deduced. Stops NPCs from magically knowing secrets or reacting to things they could not have seen.

<NOCLICHES>
Kills the cringe. No more "orbs" for eyes, "shivers down spines", or dramatic monologues. Fresh expressions, simple gestures, understated reactions.

<REALISTICDIALOGUE>
Messy human conversation - interruptions, filler words, trailing off, awkward pauses, talking over each other, mumbling. No perfect speeches.

  • 🧑‍🎨 OC
  • 💁 Assistant
  • #proxies
  • #guide
proxy allowed

Published chats

0

comments

Leave a comment or feedback for the creator ❤️