Freeze the model, collect activations for inputs, and train a simple classifier predicting the target feature from those activations.
We need to determine what information is encoded in a model's hidden representations.