rev-parse: compute object names in another hash algorithm - #2382
Conversation
Welcome to GitGitGadgetHi @xnox, and welcome to GitGitGadget, the GitHub App to send patch series to the Git mailing list from GitHub Pull Requests. Please make sure that either:
You can CC potential reviewers by adding a footer to the PR description with the following syntax: NOTE: DO NOT copy/paste your CC list from a previous GGG PR's description, Also, it is a good idea to review the commit messages one last time, as the Git project expects them in a quite specific form:
It is in general a good idea to await the automated test ("Checks") in this Pull Request before contributing the patches, e.g. to avoid trivial issues such as unportable code. Contributing the patchesBefore you can contribute the patches, your GitHub username needs to be added to the list of permitted users. Any already-permitted user can do that, by adding a comment to your PR of the form Both the person who commented An alternative is the channel Once on the list of permitted usernames, you can contribute the patches to the Git mailing list by adding a PR comment If you want to see what email(s) would be sent for a After you submit, GitGitGadget will respond with another comment that contains the link to the cover letter mail in the Git mailing list archive. Please make sure to monitor the discussion in that thread and to address comments and suggestions (while the comments and suggestions will be mirrored into the PR by GitGitGadget, you will still want to reply via mail). If you do not want to subscribe to the Git mailing list just to be able to respond to a mail, you can download the mbox from the Git mailing list archive (click the curl -g --user "<EMailAddress>:<Password>" \
--url "imaps://imap.gmail.com/INBOX" -T /path/to/raw.txtTo iterate on your change, i.e. send a revised patch or patch series, you will first want to (force-)push to the same branch. You probably also want to modify your Pull Request description (or title). It is a good idea to summarize the revision by adding something like this to the cover letter (read: by editing the first comment on the PR, i.e. the PR description): To send a new iteration, just add another PR comment with the contents: Need help?New contributors who want advice are encouraged to join git-mentoring@googlegroups.com, where volunteers who regularly contribute to Git are willing to answer newbie questions, give advice, or otherwise provide mentoring to interested contributors. You must join in order to post or view messages, but anyone can join. You may also be able to find help in real time in the developer IRC channel, |
"git rev-parse --output-object-format=sha256 HEAD^{tree}" could not
answer in a SHA-1 repository. And vice-versa. Without
extensions.compatObjectFormat the option was rejected outright, and with
it the answer was wrong: the extension records a mapping only for objects
written while it is enabled, so for the objects a repository already
contained the lookup missed, repo_oid_to_algop() left the object ID
untouched, and rev-parse printed the SHA-1 name with a zero exit code.
The machinery to convert object contents between algorithms is already
here. convert_tree_object() and friends rewrite the object IDs a tree,
commit or tag refers to, calling repo_oid_to_algop() for each one, so the
recursion is written; what is missing is a base case. Give
repo_oid_to_algop() one: on a lookup miss, read the object, convert its
contents, and hash the result. A blob needs no conversion at all, as its
contents are the same either way, only a rehash.
Names computed this way are remembered during the process execution,
so that a tree that reaches the same subtree by several paths pays for
it once. These are not stored permamently. In the future a side-car
cache could be added for these results, or eventual dual-format v3
could be used to lazy compute/convert and store these.
Commits are refused rather than computed. A commit names its parents, so
converting one converts the entire history behind it; that is a repository
conversion, not an answer about a single object, and recursing over it
here would be bounded only by the length of the history. A tree holding a
submodule is refused for the same reason, and now says which entry is at
fault instead of reporting an unreadable object.
Since names can now be computed, accept any algorithm git knows rather
than only the configured compatibility one, and check what
repo_oid_to_algop() returns so that a name we cannot produce is an error
instead of the storage name.
On git.git's tree at 9936c1b (3089 files, 26MB) this names the tree
in 0.26s, against 0.25s for reading and hashing that content by hand,
so close to the floor for the work involved. On linux.git tree of
1.6GiB it takes 8s to compute.
Initially, I have implemented contrib git-sha256-tree.sh script using
a temporary git repo and perform fast-export/fast-import of a single
commit and its tree. It is a lot slower due to needless work of
fast-import that is discarded. However, a contrib script carries less
maintainance and is more portable to existing installations. If there
is interest, I can publish that implementation as well. The goal is
to help with interop, and have the ability to record tree names in
either format today; to check again later if and when a given project
switches to sha256 format.
Signed-off-by: Dimitri John Ledkov <dimitri.ledkov@surgut.co.uk>
Assisted-by: Claude Opus 5 xHigh
0e8e735 to
5c83a90
Compare
"git rev-parse --output-object-format=sha256 HEAD^{tree}" could not
answer in a SHA-1 repository. And vice-versa. Without
extensions.compatObjectFormat the option was rejected outright, and with
it the answer was wrong: the extension records a mapping only for objects
written while it is enabled, so for the objects a repository already
contained the lookup missed, repo_oid_to_algop() left the object ID
untouched, and rev-parse printed the SHA-1 name with a zero exit code.
The machinery to convert object contents between algorithms is already
here. convert_tree_object() and friends rewrite the object IDs a tree,
commit or tag refers to, calling repo_oid_to_algop() for each one, so the
recursion is written; what is missing is a base case. Give
repo_oid_to_algop() one: on a lookup miss, read the object, convert its
contents, and hash the result. A blob needs no conversion at all, as its
contents are the same either way, only a rehash.
Names computed this way are remembered during the process execution,
so that a tree that reaches the same subtree by several paths pays for
it once. These are not stored permamently. In the future a side-car
cache could be added for these results, or eventual dual-format v3
could be used to lazy compute/convert and store these.
Commits are refused rather than computed. A commit names its parents, so
converting one converts the entire history behind it; that is a repository
conversion, not an answer about a single object, and recursing over it
here would be bounded only by the length of the history. A tree holding a
submodule is refused for the same reason, and now says which entry is at
fault instead of reporting an unreadable object.
Since names can now be computed, accept any algorithm git knows rather
than only the configured compatibility one, and check what
repo_oid_to_algop() returns so that a name we cannot produce is an error
instead of the storage name.
On git.git's tree at 9936c1b (3089 files, 26MB) this names the tree
in 0.26s, against 0.25s for reading and hashing that content by hand,
so close to the floor for the work involved. On linux.git tree of
1.6GiB it takes 8s to compute.
Initially, I have implemented contrib git-sha256-tree.sh script using
a temporary git repo and perform fast-export/fast-import of a single
commit and its tree. It is a lot slower due to needless work of
fast-import that is discarded. However, a contrib script carries less
maintainance and is more portable to existing installations. If there
is interest, I can publish that implementation as well. The goal is
to help with interop, and have the ability to record tree names in
either format today; to check again later if and when a given project
switches to sha256 format.
Signed-off-by: Dimitri John Ledkov dimitri.ledkov@surgut.co.uk
Assisted-by: Claude Opus 5 xHigh