Better handeling of hardcoded component in PretrainedModel.from_pretrained. #35617

princethewinner · 2025-01-10T19:00:42Z

System Info

transformers version: 4.42.0
Platform: Linux-5.15.0-125-generic-x86_64-with-glibc2.35
Python version: 3.10.12
Huggingface_hub version: 0.23.4
Safetensors version: 0.4.2
Accelerate version: 0.28.0
Accelerate config: not found
PyTorch version (GPU?): 2.2.1+cu121 (True)
Tensorflow version (GPU?): not installed (NA)
Flax version (CPU?/GPU?/TPU?): not installed (NA)
Jax version: not installed
JaxLib version: not installed
Using distributed or parallel set-up in script?: yes
Using GPU in script?: yes
GPU type: Tesla V100-SXM2-32GB

Who can help?

No response

Information

The official example scripts
My own modified scripts

Tasks

An officially supported task in the examples folder (such as GLUE/SQuAD, ...)
My own task or dataset (give details below)

Reproduction

The hardcoded component to replace the key names in the loaded model needs better handling (See snippets below). I had named a few variable as beta and gamma in my layers. the from_pretrained function was replacing these names with bias and weights thus loaded model was performing differently.

Probably, raise a warning/error if these names are part of layer names in the architecture to avoid trouble at the later stage of development.

https://github.com/huggingface/transformers/blob/15bd3e61f8d3680ca472c9314ad07584d20f7b81/src/transformers/modeling_utils.py#L4338C1-L4358C19

    @staticmethod
    def _fix_state_dict_key_on_load(key):
        """Replace legacy parameter names with their modern equivalents. E.g. beta -> bias, gamma -> weight."""

        if "beta" in key:
            return key.replace("beta", "bias")
        if "gamma" in key:
            return key.replace("gamma", "weight")

        # to avoid logging parametrized weight norm renaming
        if hasattr(nn.utils.parametrizations, "weight_norm"):
            if "weight_g" in key:
                return key.replace("weight_g", "parametrizations.weight.original0")
            if "weight_v" in key:
                return key.replace("weight_v", "parametrizations.weight.original1")
        else:
            if "parametrizations.weight.original0" in key:
                return key.replace("parametrizations.weight.original0", "weight_g")
            if "parametrizations.weight.original1" in key:
                return key.replace("parametrizations.weight.original1", "weight_v")
        return key

Expected behavior

Loading of the pre-trained model should not raise missing/unexpected layer warnings.

The text was updated successfully, but these errors were encountered:

Lala2398 · 2025-01-10T22:27:59Z

Thank you for bringing up this issue. Actually, it`s my first issue comment, I hope this will be help you.
You can use mapping based approach or import warnings :)
What do you think about these suggestion? Please give your feedback.

Rocketknight1 · 2025-01-13T14:32:52Z

Hi @princethewinner, we have a PR open to fix this at #35615!

github-actions · 2025-02-10T08:03:04Z

This issue has been automatically marked as stale because it has not had recent activity. If you think this still needs to be addressed please comment on this thread.

Please note that issues that do not follow the contributing guidelines are likely to be ignored.

princethewinner added the bug label Jan 10, 2025

github-actions bot closed this as completed Feb 18, 2025

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Better handeling of hardcoded component in PretrainedModel.from_pretrained. #35617

Better handeling of hardcoded component in PretrainedModel.from_pretrained. #35617

princethewinner commented Jan 10, 2025

Lala2398 commented Jan 10, 2025

Rocketknight1 commented Jan 13, 2025

github-actions bot commented Feb 10, 2025

Better handeling of hardcoded component in PretrainedModel.from_pretrained. #35617

Better handeling of hardcoded component in PretrainedModel.from_pretrained. #35617

Comments

princethewinner commented Jan 10, 2025

System Info

Who can help?

Information

Tasks

Reproduction

Expected behavior

Lala2398 commented Jan 10, 2025

Rocketknight1 commented Jan 13, 2025

github-actions bot commented Feb 10, 2025