You are using a model to grade another model's output. Where does that work, and where does it lie to you?
A judge is reliable for pairwise comparison against a concrete rubric and unreliable as an absolute scorer. It carries position, verbosity and self-preference biases and agrees with whatever the prompt implies, so validate it against human labels before trusting any number it gives you.