upstream/mercurial-mirror Files · hgext/largefiles/__init__.py

bdiff: replace hash algorithm...

bdiff: replace hash algorithm This patch replaces lyhash with the hash algorithm used by diffutils. The algorithm has its origins in Git commit 2e9d1410, which is all the way back from 1992. The license header in the code at that revision in GPL v2. I have not performed an extensive analysis of the distribution (and therefore buckets) of hash output. However, `hg perfbdiff` gives some clear wins. I'd like to think that if it is good enough for diffutils it is good enough for us? From the mozilla-unified repository: $ perfbdiff -m ! wall 0.053271 comb 0.060000 user 0.060000 sys 0.000000 (best of 100) ! wall 0.035827 comb 0.040000 user 0.040000 sys 0.000000 (best of 100) $ perfbdiff --alldata --count 100 ! wall 6.204277 comb 6.200000 user 6.200000 sys 0.000000 (best of 3) ! wall 4.309710 comb 4.300000 user 4.300000 sys 0.000000 (best of 3) From the hg repo: $ perfbdiff 35000 --alldata --count 1000 ! wall 0.660358 comb 0.660000 user 0.660000 sys 0.000000 (best of 15) ! wall 0.534092 comb 0.530000 user 0.530000 sys 0.000000 (best of 19) Looking at the generated assembly and statistical profiler output from the kernel level, I believe there is room to make this function even faster. Namely, we're still consuming data character by character instead of at the word level. This translates to more loop iterations and more instructions. At this juncture though, the real performance killer is that we're hashing every line. We should get a significant speedup if we change the algorithm to find the longest prefix, longest suffix, treat those as single "lines" and then only do the line splitting and hashing on the parts that are different. That will require a lot of C code, however. I'm optimistic this approach could result in a ~2x speedup.

Augie Fackler - - Load All Authors

File last commit:

r29841:d5883fd0 default


                r30318:e1d6aa0e

default

Download file

             __init__.py
        
                    140 lines
            
             | 5.3 KiB
            
                | text/x-python
            
             |
                PythonLexer
            
             / hgext / largefiles / __init__.py
          
                    History
                
                 |
                  Annotation
                 | Raw
                 |Copy content
                 |Copy permalink

      # Copyright 2009-2010 Gregory P. Ward

      # Copyright 2009-2010 Intelerad Medical Systems Incorporated

      # Copyright 2010-2011 Fog Creek Software

      # Copyright 2010-2011 Unity Technologies

      #

      # This software may be used and distributed according to the terms of the

      # GNU General Public License version 2 or any later version.

      '''track large binary files

      Large binary files tend to be not very compressible, not very

      diffable, and not at all mergeable. Such files are not handled

      efficiently by Mercurial's storage format (revlog), which is based on

      compressed binary deltas; storing large binary files as regular

      Mercurial files wastes bandwidth and disk space and increases

      Mercurial's memory usage. The largefiles extension addresses these

      problems by adding a centralized client-server layer on top of

      Mercurial: largefiles live in a *central store* out on the network

      somewhere, and you only fetch the revisions that you need when you

      need them.

      largefiles works by maintaining a "standin file" in .hglf/ for each

      largefile. The standins are small (41 bytes: an SHA-1 hash plus

      newline) and are tracked by Mercurial. Largefile revisions are

      identified by the SHA-1 hash of their contents, which is written to

      the standin. largefiles uses that revision ID to get/put largefile

      revisions from/to the central store. This saves both disk space and

      bandwidth, since you don't need to retrieve all historical revisions

      of large files when you clone or pull.

      To start a new repository or add new large binary files, just add

      --large to your :hg:`add` command. For example::

        $ dd if=/dev/urandom of=randomdata count=2000

        $ hg add --large randomdata

        $ hg commit -m "add randomdata as a largefile"

      When you push a changeset that adds/modifies largefiles to a remote

      repository, its largefile revisions will be uploaded along with it.

      Note that the remote Mercurial must also have the largefiles extension

      enabled for this to work.

      When you pull a changeset that affects largefiles from a remote

      repository, the largefiles for the changeset will by default not be

      pulled down. However, when you update to such a revision, any

      largefiles needed by that revision are downloaded and cached (if

      they have never been downloaded before). One way to pull largefiles

      when pulling is thus to use --update, which will update your working

      copy to the latest pulled revision (and thereby downloading any new

      largefiles).

      If you want to pull largefiles you don't need for update yet, then

      you can use pull with the `--lfrev` option or the :hg:`lfpull` command.

      If you know you are pulling from a non-default location and want to

      download all the largefiles that correspond to the new changesets at

      the same time, then you can pull with `--lfrev "pulled()"`.

      If you just want to ensure that you will have the largefiles needed to

      merge or rebase with new heads that you are pulling, then you can pull

      with `--lfrev "head(pulled())"` flag to pre-emptively download any largefiles

      that are new in the heads you are pulling.

      Keep in mind that network access may now be required to update to

      changesets that you have not previously updated to. The nature of the

      largefiles extension means that updating is no longer guaranteed to

      be a local-only operation.

      If you already have large files tracked by Mercurial without the

      largefiles extension, you will need to convert your repository in

      order to benefit from largefiles. This is done with the

      :hg:`lfconvert` command::

        $ hg lfconvert --size 10 oldrepo newrepo

      In repositories that already have largefiles in them, any new file

      over 10MB will automatically be added as a largefile. To change this

      threshold, set ``largefiles.minsize`` in your Mercurial config file

      to the minimum size in megabytes to track as a largefile, or use the

      --lfsize option to the add command (also in megabytes)::

        [largefiles]

        minsize = 2

        $ hg add --lfsize 2

      The ``largefiles.patterns`` config option allows you to specify a list

      of filename patterns (see :hg:`help patterns`) that should always be

      tracked as largefiles::

        [largefiles]

        patterns =

          *.jpg

          re:.*\.(png|bmp)$

          library.zip

          content/audio/*

      Files that match one of these patterns will be added as largefiles

      regardless of their size.

      The ``largefiles.minsize`` and ``largefiles.patterns`` config options

      will be ignored for any repositories not already containing a

      largefile. To add the first largefile to a repository, you must

      explicitly do so with the --large flag passed to the :hg:`add`

      command.

      '''

      from __future__ import absolute_import

      from mercurial import (

          hg,

          localrepo,

      )

      from . import (

          lfcommands,

          overrides,

          proto,

          reposetup,

          uisetup as uisetupmod,

      )

      # Note for extension authors: ONLY specify testedwith = 'ships-with-hg-core' for

      # extensions which SHIP WITH MERCURIAL. Non-mainline extensions should

      # be specifying the version(s) of Mercurial they are tested with, or

      # leave the attribute unspecified.

      testedwith = 'ships-with-hg-core'

      reposetup = reposetup.reposetup

      def featuresetup(ui, supported):

          # don't die on seeing a repo with the largefiles requirement

          supported |= set(['largefiles'])

      def uisetup(ui):

          localrepo.localrepository.featuresetupfuncs.add(featuresetup)

          hg.wirepeersetupfuncs.append(proto.wirereposetup)

          uisetupmod.uisetup(ui)

      cmdtable = lfcommands.cmdtable

      revsetpredicate = overrides.revsetpredicate

	Site-wide shortcuts
/	Use quick search box
g h	Goto home page
g g	Goto my private gists page
g G	Goto my public gists page
g 0-9	Goto bookmarked items from 0-9
n r	New repository page
n g	New gist page

	Repositories
g s	Goto summary page
g c	Goto changelog page
g f	Goto files page
g F	Goto files page with file search activated
g p	Goto pull requests page
g o	Goto repository settings
g O	Goto repository access permissions settings
t s	Toggle sidebar on some pages

				# Copyright 2009-2010 Gregory P. Ward
				# Copyright 2009-2010 Intelerad Medical Systems Incorporated
				# Copyright 2010-2011 Fog Creek Software
				# Copyright 2010-2011 Unity Technologies
				#
				# This software may be used and distributed according to the terms of the
				# GNU General Public License version 2 or any later version.

				'''track large binary files

				Large binary files tend to be not very compressible, not very
				diffable, and not at all mergeable. Such files are not handled
				efficiently by Mercurial's storage format (revlog), which is based on
				compressed binary deltas; storing large binary files as regular
				Mercurial files wastes bandwidth and disk space and increases
				Mercurial's memory usage. The largefiles extension addresses these
				problems by adding a centralized client-server layer on top of
				Mercurial: largefiles live in a central store out on the network
				somewhere, and you only fetch the revisions that you need when you
				need them.

				largefiles works by maintaining a "standin file" in .hglf/ for each
				largefile. The standins are small (41 bytes: an SHA-1 hash plus
				newline) and are tracked by Mercurial. Largefile revisions are
				identified by the SHA-1 hash of their contents, which is written to
				the standin. largefiles uses that revision ID to get/put largefile
				revisions from/to the central store. This saves both disk space and
				bandwidth, since you don't need to retrieve all historical revisions
				of large files when you clone or pull.

				To start a new repository or add new large binary files, just add
				--large to your :hg:`add` command. For example::

				$ dd if=/dev/urandom of=randomdata count=2000
				$ hg add --large randomdata
				$ hg commit -m "add randomdata as a largefile"

				When you push a changeset that adds/modifies largefiles to a remote
				repository, its largefile revisions will be uploaded along with it.
				Note that the remote Mercurial must also have the largefiles extension
				enabled for this to work.

				When you pull a changeset that affects largefiles from a remote
				repository, the largefiles for the changeset will by default not be
				pulled down. However, when you update to such a revision, any
				largefiles needed by that revision are downloaded and cached (if
				they have never been downloaded before). One way to pull largefiles
				when pulling is thus to use --update, which will update your working
				copy to the latest pulled revision (and thereby downloading any new
				largefiles).

				If you want to pull largefiles you don't need for update yet, then
				you can use pull with the `--lfrev` option or the :hg:`lfpull` command.

				If you know you are pulling from a non-default location and want to
				download all the largefiles that correspond to the new changesets at
				the same time, then you can pull with `--lfrev "pulled()"`.

				If you just want to ensure that you will have the largefiles needed to
				merge or rebase with new heads that you are pulling, then you can pull
				with `--lfrev "head(pulled())"` flag to pre-emptively download any largefiles
				that are new in the heads you are pulling.

				Keep in mind that network access may now be required to update to
				changesets that you have not previously updated to. The nature of the
				largefiles extension means that updating is no longer guaranteed to
				be a local-only operation.

				If you already have large files tracked by Mercurial without the
				largefiles extension, you will need to convert your repository in
				order to benefit from largefiles. This is done with the
				:hg:`lfconvert` command::

				$ hg lfconvert --size 10 oldrepo newrepo

				In repositories that already have largefiles in them, any new file
				over 10MB will automatically be added as a largefile. To change this
				threshold, set ``largefiles.minsize`` in your Mercurial config file
				to the minimum size in megabytes to track as a largefile, or use the
				--lfsize option to the add command (also in megabytes)::

				[largefiles]
				minsize = 2

				$ hg add --lfsize 2

				The ``largefiles.patterns`` config option allows you to specify a list
				of filename patterns (see :hg:`help patterns`) that should always be
				tracked as largefiles::

				[largefiles]
				patterns =
				*.jpg
				re:.*\.(png\|bmp)$
				library.zip
				content/audio/*

				Files that match one of these patterns will be added as largefiles
				regardless of their size.

				The ``largefiles.minsize`` and ``largefiles.patterns`` config options
				will be ignored for any repositories not already containing a
				largefile. To add the first largefile to a repository, you must
				explicitly do so with the --large flag passed to the :hg:`add`
				command.
				'''
				from __future__ import absolute_import

				from mercurial import (
				hg,
				localrepo,
				)

				from . import (
				lfcommands,
				overrides,
				proto,
				reposetup,
				uisetup as uisetupmod,
				)

				# Note for extension authors: ONLY specify testedwith = 'ships-with-hg-core' for
				# extensions which SHIP WITH MERCURIAL. Non-mainline extensions should
				# be specifying the version(s) of Mercurial they are tested with, or
				# leave the attribute unspecified.
				testedwith = 'ships-with-hg-core'

				reposetup = reposetup.reposetup

				def featuresetup(ui, supported):
				# don't die on seeing a repo with the largefiles requirement
				supported \|= set(['largefiles'])

				def uisetup(ui):
				localrepo.localrepository.featuresetupfuncs.add(featuresetup)
				hg.wirepeersetupfuncs.append(proto.wirereposetup)
				uisetupmod.uisetup(ui)

				cmdtable = lfcommands.cmdtable
				revsetpredicate = overrides.revsetpredicate