Skip to content

rev-parse: compute object names in another hash algorithm - #2382

Draft
xnox wants to merge 1 commit into
git:masterfrom
xnox:dl/rev-parse-output-object-format
Draft

rev-parse: compute object names in another hash algorithm#2382
xnox wants to merge 1 commit into
git:masterfrom
xnox:dl/rev-parse-output-object-format

Conversation

@xnox

@xnox xnox commented Aug 13, 2026

Copy link
Copy Markdown

"git rev-parse --output-object-format=sha256 HEAD^{tree}" could not
answer in a SHA-1 repository. And vice-versa. Without
extensions.compatObjectFormat the option was rejected outright, and with
it the answer was wrong: the extension records a mapping only for objects
written while it is enabled, so for the objects a repository already
contained the lookup missed, repo_oid_to_algop() left the object ID
untouched, and rev-parse printed the SHA-1 name with a zero exit code.

The machinery to convert object contents between algorithms is already
here. convert_tree_object() and friends rewrite the object IDs a tree,
commit or tag refers to, calling repo_oid_to_algop() for each one, so the
recursion is written; what is missing is a base case. Give
repo_oid_to_algop() one: on a lookup miss, read the object, convert its
contents, and hash the result. A blob needs no conversion at all, as its
contents are the same either way, only a rehash.

Names computed this way are remembered during the process execution,
so that a tree that reaches the same subtree by several paths pays for
it once. These are not stored permamently. In the future a side-car
cache could be added for these results, or eventual dual-format v3
could be used to lazy compute/convert and store these.

Commits are refused rather than computed. A commit names its parents, so
converting one converts the entire history behind it; that is a repository
conversion, not an answer about a single object, and recursing over it
here would be bounded only by the length of the history. A tree holding a
submodule is refused for the same reason, and now says which entry is at
fault instead of reporting an unreadable object.

Since names can now be computed, accept any algorithm git knows rather
than only the configured compatibility one, and check what
repo_oid_to_algop() returns so that a name we cannot produce is an error
instead of the storage name.

On git.git's tree at 9936c1b (3089 files, 26MB) this names the tree
in 0.26s, against 0.25s for reading and hashing that content by hand,
so close to the floor for the work involved. On linux.git tree of
1.6GiB it takes 8s to compute.

Initially, I have implemented contrib git-sha256-tree.sh script using
a temporary git repo and perform fast-export/fast-import of a single
commit and its tree. It is a lot slower due to needless work of
fast-import that is discarded. However, a contrib script carries less
maintainance and is more portable to existing installations. If there
is interest, I can publish that implementation as well. The goal is
to help with interop, and have the ability to record tree names in
either format today; to check again later if and when a given project
switches to sha256 format.

Signed-off-by: Dimitri John Ledkov dimitri.ledkov@surgut.co.uk
Assisted-by: Claude Opus 5 xHigh

@gitgitgadget-git

Copy link
Copy Markdown

Welcome to GitGitGadget

Hi @xnox, and welcome to GitGitGadget, the GitHub App to send patch series to the Git mailing list from GitHub Pull Requests.

Please make sure that either:

  • Your Pull Request has a good description, if it consists of multiple commits, as it will be used as cover letter.
  • Your Pull Request description is empty, if it consists of a single commit, as the commit message should be descriptive enough by itself.

You can CC potential reviewers by adding a footer to the PR description with the following syntax:

CC: Revi Ewer <revi.ewer@example.com>, Ill Takalook <ill.takalook@example.net>

NOTE: DO NOT copy/paste your CC list from a previous GGG PR's description,
because it will result in a malformed CC list on the mailing list. See
example.

Also, it is a good idea to review the commit messages one last time, as the Git project expects them in a quite specific form:

  • the lines should not exceed 76 columns,
  • the first line should be like a header and typically start with a prefix like "tests:" or "revisions:" to state which subsystem the change is about, and
  • the commit messages' body should be describing the "why?" of the change.
  • Finally, the commit messages should end in a Signed-off-by: line matching the commits' author.

It is in general a good idea to await the automated test ("Checks") in this Pull Request before contributing the patches, e.g. to avoid trivial issues such as unportable code.

Contributing the patches

Before you can contribute the patches, your GitHub username needs to be added to the list of permitted users. Any already-permitted user can do that, by adding a comment to your PR of the form /allow. A good way to find other contributors is to locate recent pull requests where someone has been /allowed:

Both the person who commented /allow and the PR author are able to /allow you.

An alternative is the channel #git-devel on the Libera Chat IRC network:

<newcontributor> I've just created my first PR, could someone please /allow me? https://github.com/gitgitgadget/git/pull/12345
<veteran> newcontributor: it is done
<newcontributor> thanks!

Once on the list of permitted usernames, you can contribute the patches to the Git mailing list by adding a PR comment /submit.

If you want to see what email(s) would be sent for a /submit request, add a PR comment /preview to have the email(s) sent to you. You must have a public GitHub email address for this. Note that any reviewers CC'd via the list in the PR description will not actually be sent emails.

After you submit, GitGitGadget will respond with another comment that contains the link to the cover letter mail in the Git mailing list archive. Please make sure to monitor the discussion in that thread and to address comments and suggestions (while the comments and suggestions will be mirrored into the PR by GitGitGadget, you will still want to reply via mail).

If you do not want to subscribe to the Git mailing list just to be able to respond to a mail, you can download the mbox from the Git mailing list archive (click the (raw) link), then import it into your mail program. If you use GMail, you can do this via:

curl -g --user "<EMailAddress>:<Password>" \
    --url "imaps://imap.gmail.com/INBOX" -T /path/to/raw.txt

To iterate on your change, i.e. send a revised patch or patch series, you will first want to (force-)push to the same branch. You probably also want to modify your Pull Request description (or title). It is a good idea to summarize the revision by adding something like this to the cover letter (read: by editing the first comment on the PR, i.e. the PR description):

Changes since v1:
- Fixed a typo in the commit message (found by ...)
- Added a code comment to ... as suggested by ...
...

To send a new iteration, just add another PR comment with the contents: /submit.

Need help?

New contributors who want advice are encouraged to join git-mentoring@googlegroups.com, where volunteers who regularly contribute to Git are willing to answer newbie questions, give advice, or otherwise provide mentoring to interested contributors. You must join in order to post or view messages, but anyone can join.

You may also be able to find help in real time in the developer IRC channel, #git-devel on Libera Chat. Remember that IRC does not support offline messaging, so if you send someone a private message and log out, they cannot respond to you. The scrollback of #git-devel is archived, though.

"git rev-parse --output-object-format=sha256 HEAD^{tree}" could not
answer in a SHA-1 repository.  And vice-versa.  Without
extensions.compatObjectFormat the option was rejected outright, and with
it the answer was wrong: the extension records a mapping only for objects
written while it is enabled, so for the objects a repository already
contained the lookup missed, repo_oid_to_algop() left the object ID
untouched, and rev-parse printed the SHA-1 name with a zero exit code.

The machinery to convert object contents between algorithms is already
here.  convert_tree_object() and friends rewrite the object IDs a tree,
commit or tag refers to, calling repo_oid_to_algop() for each one, so the
recursion is written; what is missing is a base case.  Give
repo_oid_to_algop() one: on a lookup miss, read the object, convert its
contents, and hash the result.  A blob needs no conversion at all, as its
contents are the same either way, only a rehash.

Names computed this way are remembered during the process execution,
so that a tree that reaches the same subtree by several paths pays for
it once.  These are not stored permamently.  In the future a side-car
cache could be added for these results, or eventual dual-format v3
could be used to lazy compute/convert and store these.

Commits are refused rather than computed.  A commit names its parents, so
converting one converts the entire history behind it; that is a repository
conversion, not an answer about a single object, and recursing over it
here would be bounded only by the length of the history.  A tree holding a
submodule is refused for the same reason, and now says which entry is at
fault instead of reporting an unreadable object.

Since names can now be computed, accept any algorithm git knows rather
than only the configured compatibility one, and check what
repo_oid_to_algop() returns so that a name we cannot produce is an error
instead of the storage name.

On git.git's tree at 9936c1b (3089 files, 26MB) this names the tree
in 0.26s, against 0.25s for reading and hashing that content by hand,
so close to the floor for the work involved. On linux.git tree of
1.6GiB it takes 8s to compute.

Initially, I have implemented contrib git-sha256-tree.sh script using
a temporary git repo and perform fast-export/fast-import of a single
commit and its tree.  It is a lot slower due to needless work of
fast-import that is discarded.  However, a contrib script carries less
maintainance and is more portable to existing installations.  If there
is interest, I can publish that implementation as well.  The goal is
to help with interop, and have the ability to record tree names in
either format today; to check again later if and when a given project
switches to sha256 format.

Signed-off-by: Dimitri John Ledkov <dimitri.ledkov@surgut.co.uk>
Assisted-by: Claude Opus 5 xHigh
@xnox
xnox force-pushed the dl/rev-parse-output-object-format branch from 0e8e735 to 5c83a90 Compare August 13, 2026 02:14
@xnox

xnox commented Aug 14, 2026

Copy link
Copy Markdown
Author

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant