upstream/mercurial-mirror Files · mercurial/similar.py

mergetools: add new conflict marker format with diffs in...

mergetools: add new conflict marker format with diffs in I use 3-way conflict markers. Often when I resolve them, I manually compare one the base with one side and apply the differences to the other side. That can be hard when the conflict marker is large. This patch introduces a new type of conflict marker, which I'm hoping will make it easier to resolve conflicts. The new format uses `<<<<<<<` and `>>>>>>>` to open and close the markers, just like our existing 2-way and 3-way conflict markers. Instead of having 2 or 3 snapshots (left+right or left+base+right), it has a sequence of diffs. A diff looks like this: ``` ------- base +++++++ left a -b +c d ``` A diff that adds one side ("diff from nothing") has a `=======` header instead and does not have have `+` prefixed on its lines. A regular 3-way merge can be viewed as adding one side plus a diff between the base and the other side. It thus has two ways of being represented, depending on which side is being diffed: ``` <<<<<<< ======= left contents on left ------- base +++++++ right contents on -left +right >>>>>>> ``` or ``` <<<<<<< ------- base +++++++ left contents on -right +left ======= right contents on right >>>>>>> ``` I've made it so the new merge tool tries to pick a version that has the most common lines (no difference in the example above). I've called the new tool "mergediff" to stick to the convention of starting with "merge" if the tool tries a regular 3-way merge. The idea came from my pet VCS (placeholder name `jj`), which has support for octopus merges and other ways of ending up with merges of more than 3 versions. I wanted to be able to represent such conflicts in the working copy and therefore thought of this format (although I have not yet implemented it in my VCS). I then attended a meeting with Larry McVoy, who said BitKeeper has an option (`bk smerge -g`) for showing a similar format, which reminded me to actually attempt this in Mercurial. Differential Revision: https://phab.mercurial-scm.org/D9551

Augie Fackler - - Load All Authors

File last commit:

r46554:89a2afe3 default


                r46724:bdc2bf68

default

Download file

             similar.py
        
                    133 lines
            
             | 4.0 KiB
            
                | text/x-python
            
             |
                PythonLexer
            
             / mercurial / similar.py
          
                    History
                
                 |
                  Source
                 | Raw
                 |Copy content
                 |Copy permalink

        David Greenaway
    
Move 'findrenames' code into its own file....

              r11059
            
      # similar.py - mechanisms for finding similar files

      #

      # Copyright 2005-2007 Matt Mackall <mpm@selenic.com>

      #

      # This software may be used and distributed according to the terms of the

      # GNU General Public License version 2 or any later version.

        Gregory Szorc
    
similar: use absolute_import

              r27359
            
      from __future__ import absolute_import

      from .i18n import _

        Gregory Szorc
    
py3: finish porting iteritems() to pycompat and remove source transformer...

              r43376
            
      from . import (

          mdiff,

          pycompat,

      )

        Augie Fackler
    
formatting: blacken the codebase...

              r43346
            
        David Greenaway
    
Move 'findrenames' code into its own file....

              r11059
            
        David Greenaway
    
findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....

              r11060
            
      def _findexactmatches(repo, added, removed):

        Augie Fackler
    
formating: upgrade to black 20.8b1...

              r46554
            
          """find renamed files that have no changes

        David Greenaway
    
findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....

              r11060
            
          Takes a list of new filectxs and a list of removed filectxs, and yields

          (before, after) tuples of exact matches.

        Augie Fackler
    
formating: upgrade to black 20.8b1...

              r46554
            
          """

        Yuya Nishihara
    
similar: use cheaper hash() function to test exact matches...

              r31584
            
          # Build table of removed files: {hash(fctx.data()): [fctx, ...]}.

          # We use hash() to discard fctx.data() from memory.

        David Greenaway
    
findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....

              r11060
            
          hashes = {}

        Augie Fackler
    
formatting: blacken the codebase...

              r43346
            
          progress = repo.ui.makeprogress(

        Augie Fackler
    
formatting: byteify all mercurial/ and hgext/ string literals...

              r43347
            
              _(b'searching for exact renames'),

        Augie Fackler
    
formatting: blacken the codebase...

              r43346
            
              total=(len(added) + len(removed)),

        Augie Fackler
    
formatting: byteify all mercurial/ and hgext/ string literals...

              r43347
            
              unit=_(b'files'),

        Augie Fackler
    
formatting: blacken the codebase...

              r43346
            
          )

        Martin von Zweigbergk
    
similar: use progress helper...

              r38367
            
          for fctx in removed:

              progress.increment()

        Yuya Nishihara
    
similar: use cheaper hash() function to test exact matches...

              r31584
            
              h = hash(fctx.data())

              if h not in hashes:

                  hashes[h] = [fctx]

              else:

                  hashes[h].append(fctx)

        David Greenaway
    
findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....

              r11060
            
          # For each added file, see if it corresponds to a removed file.

        Martin von Zweigbergk
    
similar: use progress helper...

              r38367
            
          for fctx in added:

              progress.increment()

        FUJIWARA Katsunori
    
similar: compare between actual file contents for exact identity...

              r31210
            
              adata = fctx.data()

        Yuya Nishihara
    
similar: use cheaper hash() function to test exact matches...

              r31584
            
              h = hash(adata)

              for rfctx in hashes.get(h, []):

        FUJIWARA Katsunori
    
similar: compare between actual file contents for exact identity...

              r31210
            
                  # compare between actual file contents for exact identity

                  if adata == rfctx.data():

                      yield (rfctx, fctx)

        Yuya Nishihara
    
similar: use cheaper hash() function to test exact matches...

              r31584
            
                      break

        David Greenaway
    
findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....

              r11060
            
          # Done

        Martin von Zweigbergk
    
progress: hide update(None) in a new complete() method...

              r38392
            
          progress.complete()

        David Greenaway
    
findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....

              r11060
            
        Augie Fackler
    
formatting: blacken the codebase...

              r43346
            
        Sean Farley
    
similar: move score function to module level...

              r30805
            
      def _ctxdata(fctx):

          # lazily load text

          orig = fctx.data()

          return orig, mdiff.splitnewlines(orig)

        Augie Fackler
    
formatting: blacken the codebase...

              r43346
            
        Pierre-Yves David
    
similar: remove caching from the module level...

              r30809
            
      def _score(fctx, otherdata):

          orig, lines = otherdata

          text = fctx.data()

        Yuya Nishihara
    
bdiff: proxy through mdiff module...

              r32201
            
          # mdiff.blocks() returns blocks of matching lines

        Sean Farley
    
similar: move score function to module level...

              r30805
            
          # count the number of bytes in each

          equal = 0

        Yuya Nishihara
    
bdiff: proxy through mdiff module...

              r32201
            
          matches = mdiff.blocks(text, orig)

        Sean Farley
    
similar: move score function to module level...

              r30805
            
          for x1, x2, y1, y2 in matches:

              for line in lines[y1:y2]:

                  equal += len(line)

          lengths = len(text) + len(orig)

          return equal * 2.0 / lengths

        Augie Fackler
    
formatting: blacken the codebase...

              r43346
            
        Pierre-Yves David
    
similar: remove caching from the module level...

              r30809
            
      def score(fctx1, fctx2):

          return _score(fctx1, _ctxdata(fctx2))

        Augie Fackler
    
formatting: blacken the codebase...

              r43346
            
        David Greenaway
    
findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....

              r11060
            
      def _findsimilarmatches(repo, added, removed, threshold):

        Augie Fackler
    
formating: upgrade to black 20.8b1...

              r46554
            
          """find potentially renamed files based on similar file content

        David Greenaway
    
findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....

              r11060
            
          Takes a list of new filectxs and a list of removed filectxs, and yields

          (before, after, score) tuples of partial matches.

        Augie Fackler
    
formating: upgrade to black 20.8b1...

              r46554
            
          """

        David Greenaway
    
Move 'findrenames' code into its own file....

              r11059
            
          copies = {}

        Augie Fackler
    
formatting: blacken the codebase...

              r43346
            
          progress = repo.ui.makeprogress(

        Augie Fackler
    
formatting: byteify all mercurial/ and hgext/ string literals...

              r43347
            
              _(b'searching for similar files'), unit=_(b'files'), total=len(removed)

        Augie Fackler
    
formatting: blacken the codebase...

              r43346
            
          )

        Martin von Zweigbergk
    
similar: use progress helper...

              r38414
            
          for r in removed:

              progress.increment()

        Pierre-Yves David
    
similar: remove caching from the module level...

              r30809
            
              data = None

        David Greenaway
    
Move 'findrenames' code into its own file....

              r11059
            
              for a in added:

                  bestscore = copies.get(a, (None, threshold))[1]

        Pierre-Yves David
    
similar: remove caching from the module level...

              r30809
            
                  if data is None:

                      data = _ctxdata(r)

                  myscore = _score(a, data)

        Yuya Nishihara
    
similar: take the first match instead of the last...

              r31583
            
                  if myscore > bestscore:

        David Greenaway
    
Move 'findrenames' code into its own file....

              r11059
            
                      copies[a] = (r, myscore)

        Martin von Zweigbergk
    
similar: use progress helper...

              r38414
            
          progress.complete()

        David Greenaway
    
Move 'findrenames' code into its own file....

              r11059
            
        Gregory Szorc
    
py3: finish porting iteritems() to pycompat and remove source transformer...

              r43376
            
          for dest, v in pycompat.iteritems(copies):

        Sean Farley
    
similar: rename local variable to not collide with previous...

              r30791
            
              source, bscore = v

              yield source, dest, bscore

        David Greenaway
    
Move 'findrenames' code into its own file....

              r11059
            
        Augie Fackler
    
formatting: blacken the codebase...

              r43346
            
        Yuya Nishihara
    
similar: do not look up and create filectx more than once...

              r31582
            
      def _dropempty(fctxs):

          return [x for x in fctxs if x.size() > 0]

        Augie Fackler
    
formatting: blacken the codebase...

              r43346
            
        David Greenaway
    
findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....

              r11060
            
      def findrenames(repo, added, removed, threshold):

          '''find renamed files -- yields (before, after, score) tuples'''

        Yuya Nishihara
    
similar: use common names for changectx variables...

              r31581
            
          wctx = repo[None]

          pctx = wctx.p1()

        David Greenaway
    
Move 'findrenames' code into its own file....

              r11059
            
        David Greenaway
    
findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....

              r11060
            
          # Zero length files will be frequently unrelated to each other, and

          # tracking the deletion/addition of such a file will probably cause more

          # harm than good. We strip them out here to avoid matching them later on.

        Yuya Nishihara
    
similar: do not look up and create filectx more than once...

              r31582
            
          addedfiles = _dropempty(wctx[fp] for fp in sorted(added))

          removedfiles = _dropempty(pctx[fp] for fp in sorted(removed) if fp in pctx)

        David Greenaway
    
findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....

              r11060
            
          # Find exact matches.

        Yuya Nishihara
    
similar: get rid of quadratic addedfiles.remove()...

              r31580
            
          matchedfiles = set()

          for (a, b) in _findexactmatches(repo, addedfiles, removedfiles):

              matchedfiles.add(b)

        David Greenaway
    
findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....

              r11060
            
              yield (a.path(), b.path(), 1.0)

          # If the user requested similar files to be matched, search for them also.

          if threshold < 1.0:

        Yuya Nishihara
    
similar: get rid of quadratic addedfiles.remove()...

              r31580
            
              addedfiles = [x for x in addedfiles if x not in matchedfiles]

        Augie Fackler
    
formatting: blacken the codebase...

              r43346
            
              for (a, b, score) in _findsimilarmatches(

                  repo, addedfiles, removedfiles, threshold

              ):

        David Greenaway
    
findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....

              r11060
            
                  yield (a.path(), b.path(), score)

	Site-wide shortcuts
/	Use quick search box
g h	Goto home page
g g	Goto my private gists page
g G	Goto my public gists page
g 0-9	Goto bookmarked items from 0-9
n r	New repository page
n g	New gist page

	Repositories
g s	Goto summary page
g c	Goto changelog page
g f	Goto files page
g F	Goto files page with file search activated
g p	Goto pull requests page
g o	Goto repository settings
g O	Goto repository access permissions settings
t s	Toggle sidebar on some pages

David Greenaway Move 'findrenames' code into its own file....	r11059	# similar.py - mechanisms for finding similar files
		#
		# Copyright 2005-2007 Matt Mackall <mpm@selenic.com>
		#
		# This software may be used and distributed according to the terms of the
		# GNU General Public License version 2 or any later version.

Gregory Szorc similar: use absolute_import	r27359	from __future__ import absolute_import

		from .i18n import _
Gregory Szorc py3: finish porting iteritems() to pycompat and remove source transformer...	r43376	from . import (
		mdiff,
		pycompat,
		)
Augie Fackler formatting: blacken the codebase...	r43346
David Greenaway Move 'findrenames' code into its own file....	r11059
David Greenaway findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....	r11060	def _findexactmatches(repo, added, removed):
Augie Fackler formating: upgrade to black 20.8b1...	r46554	"""find renamed files that have no changes
David Greenaway findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....	r11060
		Takes a list of new filectxs and a list of removed filectxs, and yields
		(before, after) tuples of exact matches.
Augie Fackler formating: upgrade to black 20.8b1...	r46554	"""
Yuya Nishihara similar: use cheaper hash() function to test exact matches...	r31584	# Build table of removed files: {hash(fctx.data()): [fctx, ...]}.
		# We use hash() to discard fctx.data() from memory.
David Greenaway findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....	r11060	hashes = {}
Augie Fackler formatting: blacken the codebase...	r43346	progress = repo.ui.makeprogress(
Augie Fackler formatting: byteify all mercurial/ and hgext/ string literals...	r43347	_(b'searching for exact renames'),
Augie Fackler formatting: blacken the codebase...	r43346	total=(len(added) + len(removed)),
Augie Fackler formatting: byteify all mercurial/ and hgext/ string literals...	r43347	unit=_(b'files'),
Augie Fackler formatting: blacken the codebase...	r43346	)
Martin von Zweigbergk similar: use progress helper...	r38367	for fctx in removed:
		progress.increment()
Yuya Nishihara similar: use cheaper hash() function to test exact matches...	r31584	h = hash(fctx.data())
		if h not in hashes:
		hashes[h] = [fctx]
		else:
		hashes[h].append(fctx)
David Greenaway findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....	r11060
		# For each added file, see if it corresponds to a removed file.
Martin von Zweigbergk similar: use progress helper...	r38367	for fctx in added:
		progress.increment()
FUJIWARA Katsunori similar: compare between actual file contents for exact identity...	r31210	adata = fctx.data()
Yuya Nishihara similar: use cheaper hash() function to test exact matches...	r31584	h = hash(adata)
		for rfctx in hashes.get(h, []):
FUJIWARA Katsunori similar: compare between actual file contents for exact identity...	r31210	# compare between actual file contents for exact identity
		if adata == rfctx.data():
		yield (rfctx, fctx)
Yuya Nishihara similar: use cheaper hash() function to test exact matches...	r31584	break
David Greenaway findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....	r11060
		# Done
Martin von Zweigbergk progress: hide update(None) in a new complete() method...	r38392	progress.complete()
David Greenaway findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....	r11060
Augie Fackler formatting: blacken the codebase...	r43346
Sean Farley similar: move score function to module level...	r30805	def _ctxdata(fctx):
		# lazily load text
		orig = fctx.data()
		return orig, mdiff.splitnewlines(orig)

Augie Fackler formatting: blacken the codebase...	r43346
Pierre-Yves David similar: remove caching from the module level...	r30809	def _score(fctx, otherdata):
		orig, lines = otherdata
		text = fctx.data()
Yuya Nishihara bdiff: proxy through mdiff module...	r32201	# mdiff.blocks() returns blocks of matching lines
Sean Farley similar: move score function to module level...	r30805	# count the number of bytes in each
		equal = 0
Yuya Nishihara bdiff: proxy through mdiff module...	r32201	matches = mdiff.blocks(text, orig)
Sean Farley similar: move score function to module level...	r30805	for x1, x2, y1, y2 in matches:
		for line in lines[y1:y2]:
		equal += len(line)

		lengths = len(text) + len(orig)
		return equal * 2.0 / lengths

Augie Fackler formatting: blacken the codebase...	r43346
Pierre-Yves David similar: remove caching from the module level...	r30809	def score(fctx1, fctx2):
		return _score(fctx1, _ctxdata(fctx2))

Augie Fackler formatting: blacken the codebase...	r43346
David Greenaway findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....	r11060	def _findsimilarmatches(repo, added, removed, threshold):
Augie Fackler formating: upgrade to black 20.8b1...	r46554	"""find potentially renamed files based on similar file content
David Greenaway findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....	r11060
		Takes a list of new filectxs and a list of removed filectxs, and yields
		(before, after, score) tuples of partial matches.
Augie Fackler formating: upgrade to black 20.8b1...	r46554	"""
David Greenaway Move 'findrenames' code into its own file....	r11059	copies = {}
Augie Fackler formatting: blacken the codebase...	r43346	progress = repo.ui.makeprogress(
Augie Fackler formatting: byteify all mercurial/ and hgext/ string literals...	r43347	_(b'searching for similar files'), unit=_(b'files'), total=len(removed)
Augie Fackler formatting: blacken the codebase...	r43346	)
Martin von Zweigbergk similar: use progress helper...	r38414	for r in removed:
		progress.increment()
Pierre-Yves David similar: remove caching from the module level...	r30809	data = None
David Greenaway Move 'findrenames' code into its own file....	r11059	for a in added:
		bestscore = copies.get(a, (None, threshold))[1]
Pierre-Yves David similar: remove caching from the module level...	r30809	if data is None:
		data = _ctxdata(r)
		myscore = _score(a, data)
Yuya Nishihara similar: take the first match instead of the last...	r31583	if myscore > bestscore:
David Greenaway Move 'findrenames' code into its own file....	r11059	copies[a] = (r, myscore)
Martin von Zweigbergk similar: use progress helper...	r38414	progress.complete()
David Greenaway Move 'findrenames' code into its own file....	r11059
Gregory Szorc py3: finish porting iteritems() to pycompat and remove source transformer...	r43376	for dest, v in pycompat.iteritems(copies):
Sean Farley similar: rename local variable to not collide with previous...	r30791	source, bscore = v
		yield source, dest, bscore
David Greenaway Move 'findrenames' code into its own file....	r11059
Augie Fackler formatting: blacken the codebase...	r43346
Yuya Nishihara similar: do not look up and create filectx more than once...	r31582	def _dropempty(fctxs):
		return [x for x in fctxs if x.size() > 0]

Augie Fackler formatting: blacken the codebase...	r43346
David Greenaway findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....	r11060	def findrenames(repo, added, removed, threshold):
		'''find renamed files -- yields (before, after, score) tuples'''
Yuya Nishihara similar: use common names for changectx variables...	r31581	wctx = repo[None]
		pctx = wctx.p1()
David Greenaway Move 'findrenames' code into its own file....	r11059
David Greenaway findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....	r11060	# Zero length files will be frequently unrelated to each other, and
		# tracking the deletion/addition of such a file will probably cause more
		# harm than good. We strip them out here to avoid matching them later on.
Yuya Nishihara similar: do not look up and create filectx more than once...	r31582	addedfiles = _dropempty(wctx[fp] for fp in sorted(added))
		removedfiles = _dropempty(pctx[fp] for fp in sorted(removed) if fp in pctx)
David Greenaway findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....	r11060
		# Find exact matches.
Yuya Nishihara similar: get rid of quadratic addedfiles.remove()...	r31580	matchedfiles = set()
		for (a, b) in _findexactmatches(repo, addedfiles, removedfiles):
		matchedfiles.add(b)
David Greenaway findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....	r11060	yield (a.path(), b.path(), 1.0)

		# If the user requested similar files to be matched, search for them also.
		if threshold < 1.0:
Yuya Nishihara similar: get rid of quadratic addedfiles.remove()...	r31580	addedfiles = [x for x in addedfiles if x not in matchedfiles]
Augie Fackler formatting: blacken the codebase...	r43346	for (a, b, score) in _findsimilarmatches(
		repo, addedfiles, removedfiles, threshold
		):
David Greenaway findrenames: Optimise "addremove -s100" by matching files by their SHA1 hashes....	r11060	yield (a.path(), b.path(), score)